VLDB 2026 Research / reviewers in the wild / expert
Toshihiko Yamasaki
dblp:81/881
· DBLP profile ↗
174ranked-venue papers
23as first author
62since 2021 · last 2026
0000-0002-1784-2314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 153 · 19 first-author · 46 since 2021Artificial intelligence and machine learning · 51 · 6 first-author · 28 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difficulty Controlled Diffusion Model for Synthesizing Effective Training DataabstractGenerative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target dataset distribution. As a result, they capture only the common features in the real dataset and mostly generate "easy samples", which are already well learned by models trained on real data. In contrast, those rare "hard samples", with atypical features but crucial for enhancing performance, cannot be effectively generated. Consequently, these approaches must synthesize large volumes of data to yield appreciable performance gains, yet the improvement remains limited. To overcome this limitation, we present a novel method that can learn to control the learning difficulty of samples during generation while also achieving domain alignment. Thus, it can efficiently generate valuable "hard samples" that yield significant performance improvements for target tasks. This is achieved by incorporating learning difficulty as an additional conditioning signal in generative models, together with a designed encoder structure and training–generation strategy. Experimental results across multiple datasets show that our method can achieve higher performance with lower generation cost. Specifically, we obtain the best performance with only 10% additional synthetic data, saving 63.4 GPU hours of generation time compared to the previous SOTA on ImageNet. Moreover, our method provides insightful visualizations of category-specific hard factors, serving as a tool for analyzing datasets. Zerun Wang, Jiafeng Mao, Toshihiko Yamasaki |
AAAI | 4 |
| 2026 | ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing PointsabstractRecent text-to-image models, such as Stable Diffusion, have achieved impressive visual quality, yet they often suffer from geometric inconsistencies that undermine the structural realism of generated scenes. One prominent issue is vanishing point inconsistency, where projections of parallel lines fail to converge correctly in 2D space. This leads to structurally implausible geometry that degrades spatial realism, especially in architectural scenes. We propose ControlVP, a user-guided framework for correcting vanishing point inconsistencies in generated images. Our approach extends a pre-trained diffusion model by incorporating structural guidance derived from building contours. We also introduce geometric constraints that explicitly encourage alignment between image edges and perspective cues. Our method enhances global geometric consistency while maintaining visual fidelity comparable to the baselines. This capability is particularly valuable for applications that require accurate spatial structure, such as image-to-3D reconstruction. The dataset and source code are available at https://github.com/RyotaOkumura/ControlVP. Ryota Okumura, Kaede Shiohara, Toshihiko Yamasaki |
WACV | 3 |
| 2026 | Language-guided frameworks for personalized video summarizationabstractExisting video summarization methods predominantly produce generic summaries and often fail to reflect user-specific preferences. To address this limitation, we explore the potential of large language models (LLMs) for video summarization and propose two language-guided frameworks for personalized video summarization. We first propose Few-Shot Video SUMmarization (FS-VSUM), a non-trainable, example-driven framework that leverages LLM-based semantic reasoning to perform annotator-personalized video summarization. By conditioning on a small number of annotated examples, FS-VSUM captures annotator-specific summarization styles and generates customized summaries without parameter updates, demonstrating the inherent capability of LLMs for controllable and personalized video summarization. We then introduce Self-Supervised Video SUMmarization (SS-VSUM), a trainable framework that formulates video summarization as a semantic textual similarity task. SS-VSUM incorporates user preferences through LLM prompts and introduces a Preserving Diversity Loss (PDL) to dynamically regulate regularization based on linguistic diversity. We further extend SS-VSUM with additional analyses and clarifications, providing a more systematic understanding of language-guided video summarization. Experimental results show that SS-VSUM achieves state-of-the-art performance on the SumMe dataset. Together, this work provides a systematic investigation of language-guided video summarization, revealing how LLMs can support both training-free personalization and trainable performance optimization. The source code for the proposed frameworks is publicly available at https://github.com/sugitomoo/VSUM. Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
Pattern Anal. Appl. | 4 |
| 2025 | Medical Clinic Revenue Prediction Using Latent Feature Extraction from Satellite Imagery with Large Multimodal Models
Shuntaro Masuda, Fumiya Matsuno, Itsuki Hirai, Koji Muta, Toshihiko Yamasaki |
ADMA (3) | 5 |
| 2025 | Adaptive Multimodal Transformer for Personality Trait Assessment in Online Job Interviews
Shengzhou Yi, Toshiaki Yamasaki, Toshihiko Yamasaki |
ADMA (3) | 3 |
| 2025 | Robust Deepfake Detection for Electronic Know Your Customer Systems Using Registered ImagesabstractIn this paper, we present a deepfake detection algorithm specifically designed for electronic Know Your Customer (eKYC) systems. To ensure the reliability of eKYC systems against deepfake attacks, it is essential to develop a robust deepfake detector capable of identifying both face swapping and face reenactment, while also being robust to image degradation. We address these challenges through three key contributions: (1) Our approach evaluates the video’s authenticity by detecting temporal inconsistencies in identity vectors extracted by face recognition models, leading to comprehensive detection of both face swapping and face reenactment. (2) In addition to processing video input, the algorithm utilizes a registered image (assumed to be genuine) to calculate identity discrepancies between the input video and the registered image, significantly improving detection accuracy. (3) We find that employing a face feature extractor trained on a larger dataset enhances both detection performance and robustness against image degradation. Our experimental results show that our proposed method accurately detects both face swapping and face reenactment comprehensively and is robust against various forms of unseen image degradation. Our source code is publicly available https://github.com/TaikiMiyagawa/DeepfakeDetection4eKYC. Takuma Amada, Kazuya Kakizaki, Taiki Miyagawa, Akinori F. Ebihara, Kaede Shiohara, Toshihiko Yamasaki |
FG | 6 |
| 2025 | ActRecognition-GPT: Utilizing Multimodal Large Language Models for Spatiotemporal Action Recognition in Nursery VideosabstractSpatiotemporal action recognition in nursery videos is essential for intelligent childcare systems that can automatically monitor children’s behaviors, detect safety risks, and generate accurate activity logs without increasing caregivers’ burden. However, occlusions, dynamic multi-person interactions, and subtle distinctions between similar actions present significant challenges. To address these challenges, we introduce ActRecognition-GPT, a multimodal framework that combines visual object detection with the reasoning capability of multimodal large language models (MLLMs). By assigning unique person IDs and incorporating both spatial and temporal context through structured prompts, the model achieves temporally consistent recognition of individual behaviors, even under visual ambiguity or occlusion. Moreover, predefined vocabularies are used to constrain generation and reduce semantic drift. Crucially, the open-vocabulary nature of MLLMs enables ActRecognition-GPT to generalize to previously unseen or ambiguous scenarios without requiring extensive labeled data, which is an essential property for modeling the diverse and evolving behaviors of children. Experimental results on a real-world nursery video dataset demonstrate a significant improvement in mean Average Precision over single-frame baselines, validating the effectiveness of combining structured visual inputs with LLM-based temporal reasoning. Kenta Watanabe, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
FG | 4 |
| 2025 | CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
Michihiro Kuroki, Toshihiko Yamasaki |
ICCV | 2 |
| 2025 | Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-ExplanationabstractImage scoring is a crucial task in numerous real-world applications. To trust a model’s judgment, understanding its rationale is essential. This paper proposes a novel training method for Vision Language Models (VLMs) to generate not only image scores but also corresponding justifications in natural language. Leveraging only an image scoring dataset and an instruction-tuned VLM, our method enables self-training, utilizing the VLM’s generated text without relying on external data or models. In addition, we introduce a simple method for creating a dataset designed to improve alignment between predicted scores and their textual justifications. By iteratively training the model with Direct Preference Optimization on two distinct datasets and merging them, we can improve both scoring accuracy and the coherence of generated explanations. Naoto Tanji, Toshihiko Yamasaki |
ICIP | 2 |
| 2025 | TourMLLM: A Retrieval-Augmented Multimodal Large Language Model for Multitask Learning in the Tourism Domain
Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki |
ICMR | 3 |
| 2025 | LITA: LMM-Guided Image-Text Alignment for Art Assessment
Tatsumi Sunada, Kaede Shiohara, Ling Xiao 0001, Toshihiko Yamasaki |
MMM (2) | 4 |
| 2025 | Multi-level knowledge distillation for fine-grained fashion image retrieval
Ling Xiao 0001, Toshihiko Yamasaki |
Knowl. Based Syst. | 2 |
| 2024 | Face2Diffusion for Fast and Editable Face PersonalizationabstractFace personalization aims to insert specific faces, taken from images, into pretrained text-to-image diffusion mod-els. However, it is still challenging for previous meth-ods to preserve both the identity similarity and editabil-ity due to overfitting to training samples. In this pa-per, we propose Face2Diffusion (F2D) for high-editability face personalization. The core idea behind F2D is that removing identity-irrelevant information from the training pipeline prevents the overfitting problem and improves ed-itability of encoded faces. F2D consists of the following three novel components: 1) Multi-scale identity en-coder provides well-disentangled identity features while keeping the benefits of multi-scale information, which im-proves the diversity of camera poses. 2) Expression guid-ance disentangles face expressions from identities and im-proves the controllability of face expressions. 3) Class-guided denoising regularization encourages models to learn how faces should be denoised, which boosts the text-alignment of backgrounds. Extensive experiments on the FaceForensics++ dataset and diverse prompts demonstrate our method greatly improves the trade-off between the identity- and text-fidelity compared to previous state-of-the-art methods. Code is available at https://github.com/mapooon/Face2Diffusion. Kaede Shiohara, Toshihiko Yamasaki |
CVPR | 2 |
| 2024 | Improving Plasticity in Online Continual Learning via Collaborative LearningabstractOnline Continual Learning (CL) solves the problem of learning the ever-emerging new classification tasks from a continuous data stream. Unlike its offline counterpart, in online CL, the training data can only be seen once. Most existing online CL research regards catastrophic forgetting (i.e., model stability) as almost the only challenge. In this paper, we argue that the model's capability to acquire new knowledge (i.e., model plasticity) is another challenge in online CL. While replay-based strategies have been shown to be effective in alleviating catastrophic forgetting, there is a notable gap in research attention toward improving model plasticity. To this end, we propose Collaborative Continual Learning (CCL), a collaborative learning based strategy to improve the model's capability in acquiring new concepts. Additionally, we introduce Distillation Chain (DC), a collaborative learning scheme to boost the training of the models. We adapt CCL-DC to existing representative online CL works. Extensive experiments demonstrate that even if the learners are well-trained with state-of-the-art online CL methods, our strategy can still improve model plasticity dramatically, and thereby improve the overall performance by a large margin. The source code of our work is available at https://github.com/maorong-wang/CCL-DC. Maorong Wang, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki |
CVPR | 4 |
| 2024 | SCOMatch: Alleviating Overtrusting in Open-Set Semi-supervised Learning
Zerun Wang, Liuyu Xiang, Lang Huang 0001, Jiafeng Mao, Ling Xiao 0001, Toshihiko Yamasaki |
ECCV (51) | 6 |
| 2024 | Adversarial Robustness of Convolutional Models Learned in the Frequency DomainabstractThis paper presents an extensive comparison of the noise robustness of standard Convolutional Neural Networks (CNNs) trained on image inputs and those trained in the frequency domain. We investigate the robustness of CNNs to small adversarial noise in the RGB input space and show that CNNs trained on Discrete Cosine Transform (DCT) inputs exhibit significantly better noise robustness to both adversarial and common spatial transformations compared to standard CNNs learned on RGB/Grayscale input. Our results suggest that frequency-domain learning of convolutional models may disentangle frequencies corresponding to semantic and adversarial features, resulting in improved adversarial robustness. This research highlights the potential of frequency domain learning to improve neural network robustness to test-time noise and warrants further investigation in this area. Subhajit Chaudhury, Toshihiko Yamasaki |
ICASSP | 2 |
| 2024 | Explaining 3D Object Detection Through Shapley Value-Based Attribution MapabstractArtificial Intelligence (AI)-based 3D object detection utilizing point clouds from LiDAR sensors has become widespread in various applications, such as autonomous driving. However, the lack of transparency in AI decision-making can result in inaccurate detections in unknown situations, potentially posing safety risks. Although explainable AI (XAI) has recently gained attention as a method to elucidate the rationale behind AI inferences, most existing methods are designed for image-based tasks, with limited methods specifically addressing point clouds and object detection. In this study, we propose 3D-SVAM, which provides explanations for 3D object detection through Shapley Value-based Attribution Map. The Shapley value can justify explanations by adhering to desirable properties for interpretability. Despite its substantial computational complexity, our method efficiently mitigates this complexity by introducing a suitable approximation. Our method shows superior performance compared to the state-of-the-art method through quantitative evaluations. Moreover, we demonstrate applications of our method, such as analyzing detection robustness to changes in point cloud distribution and correcting false detection by identifying points that negatively contribute to the prediction. These findings clarify the properties of 3D object detection and enhance the practical application of XAI. Michihiro Kuroki, Toshihiko Yamasaki |
ICIP | 2 |
| 2024 | Adversarially Robust Continual Learning with Anti-Forgetting LossabstractExisting continual learning methods focus on preventing catastrophic forgetting but often overlook the challenge of adversarial examples in image classification. In this study, we propose a novel method that balances accuracy, robustness against adversarial examples, and the prevention of forgetting. Specifically, we first theoretically and experimentally demonstrate that learning through knowledge distillation, a common strategy in continual learning, conflicts with learning through the cross-entropy loss. To resolve this conflict, we propose a novel loss function that combines an additional memory data loss with a conflict-avoiding knowledge distillation loss, effectively preventing catastrophic forgetting while ensuring robustness. Experimental results show that the proposed method outperforms existing methods by 5.17% in clean accuracy and $2.10 \%$ in robust accuracy. This method proves to be especially beneficial in scenarios where the reuse of samples from previous tasks is limited. Koki Mukai, Soichiro Kumano, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki |
ICIP | 5 |
| 2024 | Theoretical Understanding of Learning from Adversarial PerturbationsabstractIt is not fully understood why adversarial examples can deceive neural networks and transfer between different networks. To elucidate this, several studies have hypothesized that adversarial perturbations, while appearing as noises, contain class features. This is supported by empirical evidence showing that networks trained on mislabeled adversarial examples can still generalize well to correctly labeled test samples. However, a theoretical understanding of how perturbations include class features and contribute to generalization is limited. In this study, we provide a theoretical framework for understanding learning from perturbations using a one-hidden-layer network trained on mutually orthogonal samples. Our results highlight that various adversarial perturbations, even perturbations of a few pixels, contain sufficient class features for generalization. Moreover, we reveal that the decision boundary when learning from perturbations matches that from standard samples except for specific regions under mild conditions. The code is available at https://github.com/s-kumano/learning-from-adversarial-perturbations. Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki |
ICLR | 3 |
| 2024 | Rethinking Momentum Knowledge Distillation in Online Continual LearningabstractOnline Continual Learning (OCL) addresses the problem of training neural networks on a continuous data stream where multiple classification tasks emerge in sequence. In contrast to offline Continual Learning, data can be seen only once in OCL, which is a very severe constraint. In this context, replay-based strategies have achieved impressive results and most state-of-the-art approaches heavily depend on them. While Knowledge Distillation (KD) has been extensively used in offline Continual Learning, it remains under-exploited in OCL, despite its high potential. In this paper, we analyze the challenges in applying KD to OCL and give empirical justifications. We introduce a direct yet effective methodology for applying Momentum Knowledge Distillation (MKD) to many flagship OCL methods and demonstrate its capabilities to enhance existing approaches. In addition to improving existing state-of-the-art accuracy by more than $10%$ points on ImageNet100, we shed light on MKD internal mechanics and impacts during training in OCL. We argue that similar to replay, MKD should be considered a central component of OCL. The code is available at https://github.com/Nicolas1203/mkd_ocl. Nicolas Michel, Maorong Wang, Ling Xiao 0001, Toshihiko Yamasaki |
ICML | 4 |
| 2024 | Holistic Visualization of Contextual Knowledge in Hotel Customer Reviews Using Self-AttentionabstractTraditional methods for analyzing customer reviews, such as word clouds and sentiment analysis, often fall short in providing comprehensive insights for business decision-making. Recent advancements in natural language processing (NLP) have improved interpretability, yet existing approaches struggle to simultaneously extract practically significant opinions, preserve contextual information, and offer intuitive visualizations of overall trends. This study proposes a novel visualization method for key opinion extraction that addresses these challenges by integrating self-attention models with dependency relationships. Our approach combines attention visualization with semantic relationships between words, enabling a holistic view of customer opinions while maintaining individual review context. We demonstrate the practical utility of our method using customer review data from a major Japanese hotel search site, revealing significant factors influencing both low and high hotel ratings. The results highlight our method’s effectiveness in extracting actionable insights from customer reviews, offering valuable implications for the hospitality industry and beyond. Shuntaro Masuda, Toshihiko Yamasaki |
ISM | 2 |
| 2024 | A Web Demo Interface for Super-Resolution Reconstruction with Parametric Regularization LossabstractThis paper presents a demo of our novel approach to improve single-image super-resolution methods by integrating trainable regularization techniques. Recent advancements, such as the noise Enhanced Super Resolution Generative Adversarial Network Plus (nESRGAN+), have shown promising results in enhancing the performance of ESRGAN. However, despite its success, nESRGAN+ still faces limitations in perceptual quality due to the absence of detailed hallucinations and the presence of unwanted artifacts with slow convergence rates. To address these challenges, we propose the integration of multiple parametric regularization algorithms, enabling iterative adjustment of network gradients. Through a series of experiments, we demonstrate that our approach yields high-quality reconstructed images, effectively restoring complex textures even in previously unseen scenarios. Moreover, the introduced loss functions contribute to accelerated convergence rates and substantial improvements in the visual fidelity of the reconstructed outputs. Our online demo system can accept input images and show the super-resolution images using our method and the two state-of-the-art methods. Supatta Viriyavisuthisakul, Parinya Sanguansat, Toshihiko Yamasaki |
ICMR | 3 |
| 2024 | Enhancing Speaking and Slide Design Skills with Deep Learning: An Online Presentation Assessment SystemabstractPresentation skills, which involve the effective use of verbal and nonverbacl cues, enable audiences to better understand the content being presented. We develope a deep learning-based online assessment system that can objectively evaluate speakers' oral presentations and slide design, providing comprehensive feedback to support their self-practice. For the speaking skill assessment, we construct a multimodal neural network, including LSTMs and attention networks, to analyze the linguistic and acoustic features of oral presentations. The proposed model can predict 14 distinct types of audience impressions with an average accuracy of 85.0%. For the slide design assessment, we propose a method that can analyze slide design based on their visual and structural features, independent of file formats. It can determine whether the slides meet 10 assessment criteria with an average accuracy of 81.7%. Shengzhou Yi, Junichiro Matsugami, Takuya Yamamoto, Toshihiko Yamasaki |
ACM Multimedia | 4 |
| 2024 | Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
MMAsia | 4 |
| 2024 | Wide Two-Layer Networks can Learn from Adversarial PerturbationsabstractAdversarial examples have raised several open questions, such as why they can deceive classifiers and transfer between different models. A prevailing hypothesis to explain these phenomena suggests that adversarial perturbations appear as random noise but contain class-specific features. This hypothesis is supported by the success of perturbation learning, where classifiers trained solely on adversarial examples and the corresponding incorrect labels generalize well to correctly labeled test data. Although this hypothesis and perturbation learning are effective in explaining intriguing properties of adversarial examples, their solid theoretical foundation is limited. In this study, we theoretically explain the counterintuitive success of perturbation learning. We assume wide two-layer networks and the results hold for any data distribution. We prove that adversarial perturbations contain sufficient class-specific features for networks to generalize from them. Moreover, the predictions of classifiers trained on mislabeled adversarial examples coincide with those of classifiers trained on correctly labeled clean samples. The code is available at https://github.com/s-kumano/perturbation-learning. Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki |
NeurIPS | 3 |
| 2024 | Dealing with Synthetic Data Contamination in Online Continual LearningabstractImage generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine learning community that are not clearly identified. Meanwhile, the success of deep learning in computer vision is driven by the massive dataset collected on the Internet. The extensive quantity of synthetic data being added to the Internet would become an obstacle for future researchers to collect "clean" datasets without AI-generated content. Prior research has shown that using datasets contaminated by synthetic images may result in performance degradation when used for training. In this paper, we investigate the potential impact of contaminated datasets on Online Continual Learning (CL) research. We experimentally show that contaminated datasets might hinder the training of existing online CL methods. Also, we propose Entropy Selection with Real-synthetic similarity Maximization (ESRM), a method to alleviate the performance deterioration caused by synthetic images when training online CL models. Experiments show that our method can significantly alleviate performance deterioration, especially when the contamination is severe. For reproducibility, the source code of our work is available at https://github.com/maorong-wang/ESRM. Maorong Wang, Nicolas Michel, Jiafeng Mao, Toshihiko Yamasaki |
NeurIPS | 4 |
| 2024 | E-ReaRev: Adaptive Reasoning for Question Answering over Incomplete Knowledge Graphs by Edge and Meaning Extensions
Xiaotong Ye, Ling Xiao 0001, Toshihiko Yamasaki |
NLDB (2) | 4 |
| 2024 | LLaVA-Tour: A Large Multimodal Model for Japanese Tourist Spot Prediction and Review GenerationabstractTourist landmark recognition and review generation can significantly enhance travel planning by helping travelers make informed decisions, optimize experiences, and discover new destinations. These capabilities also benefit local economies and improve business efficiency in tourism. The advancements in large multimodal models have demonstrated high performance across a wide range of image processing tasks due to their extensive knowledge and reasoning capabilities. However, there are no large-scale tourism datasets and no models specifically tailored for tourism have been developed. To address these issues, we created a new dataset with over 1.3 million entries from Japanese tourism website Jalan.net, covering three types of tasks: landmark recognition, general and conditioned review generation and one support task: description generation. Using the Large Language-and-Vision Assistant (LLaVA) as the baseline, we developed instruction tuning strategies for these tasks and created a new Large Multimodal Model, LLaVA-Tour. By integrating domain-specific knowledge, our model outperformed state-of-the-art large multimodal models and review generation models in landmark recognition, and general and conditioned review generation. The code and data url are available at https://github.com/HiromasaYamanishi/LLaVATour. Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki |
VCIP | 3 |
| 2024 | Transferability prediction among classification and regression tasks using optimal transportabstractAbstract Transfer learning is a method for improving generalization performance by training a model for a different task first and then additionally training the pre-learned weights for the target task. However, transferability—the ease with which source task can be effectively transferred to which target task—is often unknown. Existing works proposed methods of measuring the transferability between classification tasks using images and discrete labels, but it cannot be applied to regression tasks. In this work, we investigate transferability among classification and regression tasks, and propose a method for predicting transferability by extending the optimal transport theory. Our transferability prediction model also can be applied to subjective tasks (e.g., aesthetics and memorability), which are usually regression tasks. We show that the appropriate source (pre-training) tasks can be predicted for the chosen target task without conducting actual pre-training and transferring trials. Experimental results demonstrated high prediction accuracy (correlation coefficient of $$\rho =0.791$$ ρ = 0.791 ) and a speed improvement of approximately 300 times compared with the above-mentioned greedy approach. Tomoyuki Hatakeyama, Toshihiko Yamasaki |
Multim. Tools Appl. | 3 |
| 2024 | A large-scale television advertising dataset for detailed impression analysisabstractAbstract Creating impressive video content such as movies and advertisements is a very important yet challenging task in business that requires both a sense of creativity and a lot of experience. Even professionals cannot necessarily invoke the impressions and emotions that they have aimed at. Many video advertisements are created and then disappear without giving a large impact on viewers. This paper presents a large-scale dataset of television (TV) advertisements that consists of 14,490 videos. The impressions of each video such as the recognition rate and interestingness rate are from the results of questionnaires answered by 620 participants. We also present a baseline for predicting the impression effects of TV advertisements by using visual and audio information, metadata such as broadcasting pattern, business category, the popularity of the casts, and text information including texts appearing on videos and narrations in audios. We predict four impressions of the viewers: 1) how much participants remember the video afterward, 2) how much they feel like buying the product/service, 3) how much they become interested in the product/service, and 4) how much they like the content of the advertisement itself. By combining images, audio, metadata, cast data, and text data, our baseline method is able to predict such impressions with a correlation of 0.69-0.82, much better than using a single-modal feature such as visual data or audio data only. This paper also gives some possible applications such as estimating the importance scores of each key frame, which gives us informative insights about how to make the advertisement content more impressive. Shunsuke Nakamura, Tatsuya Kawahara, Gen Tamura, Toshihiko Yamasaki |
Multim. Tools Appl. | 6 |
| 2024 | Personalized Image Enhancement Featuring Masked Style ModelingabstractWe address personalized image enhancement in this study, where we enhance input images for each user based on the user’s preferred images. Previous methods apply the same preferred style to all input images (i.e., only one style for each user); in contrast to these methods, we aim to achieve content-aware personalization by applying different styles to each image considering the contents. For content-aware personalization, we make two contributions. First, we propose a method named masked style modeling, which can predict a style for an input image considering the contents by using the framework of masked language modeling. Second, to allow this model to consider the contents of images, we propose a novel training scheme where we download images from Flickr and create pseudo input and retouched image pairs using a degrading model. We conduct quantitative evaluations and a user study, and our method trained using our training scheme successfully achieves content-aware personalization; moreover, our method outperforms other previous methods in this field. Our source code is available athttps://github.com/satoshi-kosugi/masked-style-modeling. Satoshi Kosugi, Toshihiko Yamasaki |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | An Online Presentation Slide Assessment System Using Visual and Semantic Segmentation Features
Shengzhou Yi, Junichiro Matsugami, Hiroshi Yumoto, Toshihiko Yamasaki |
AAAI | 4 |
| 2023 | Adversarial Training from Mean Field PerspectiveabstractAlthough adversarial training is known to be effective against adversarial examples, training dynamics are not well understood. In this study, we present the first theoretical analysis of adversarial training in random deep neural networks without any assumptions on data distributions. We introduce a new theoretical framework based on mean field theory, which addresses the limitations of existing mean field-based approaches. Based on the framework, we derive the (empirically tight) upper bounds of $\ell_q$ norm-based adversarial loss with $\ell_p$ norm-based adversarial examples for various values of $p$ and $q$. Moreover, we prove that networks without shortcuts are generally not adversarially trainable and that adversarial training reduces network capacity. We also show that the network width alleviates these issues. Furthermore, the various impacts of input and output dimensions on the upper bounds and time evolution of weight variance are presented. Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki |
NeurIPS | 3 |
| 2023 | Image aesthetics prediction using multiple patches preserving the original aspect ratio of contents
Toshihiko Yamasaki |
Multim. Tools Appl. | 3 |
| 2023 | Which account will you follow? Recommending influential accounts on social media
Yiwei Zhang 0014, Toshihiko Yamasaki |
Multim. Tools Appl. | 3 |
| 2023 | Parametric loss-based super-resolution for scene text recognition
Supatta Viriyavisuthisakul, Parinya Sanguansat, Teeradaj Racharak, Minh Le Nguyen 0001, Natsuda Kaothanthong, Choochart Haruechaiyasak, Toshihiko Yamasaki |
Mach. Vis. Appl. | 7 |
| 2023 | Sparse fooling images: Fooling machine perception through unrecognizable imagesabstract• Revealing a new vulnerability of DNNs through SFIs. • SFIs neither have features, and they distribute extremely far from natural images. • Proving the existence of SFIs under mild conditions for three models. • Theoretically indicating that complex models are more vulnerable to SFIs. • Experimentally confirming the threat by SFI for various datasets and models. Fooling images are potential threats to deep neural networks (DNNs). These images cannot be recognized by humans as natural objects, e.g., dogs and cats. However, they are misclassified by DNNs as natural object classes with high confidence scores. Despite their original design concept, existing fooling images, if closely examined, can be seen to retain some features that are characteristic of the target objects. Hence, DNNs can react to these features. In this study, we evaluate whether fooling images with no characteristic pattern of natural objects, either locally or globally, can exist. As a minimal case, we introduce single-color images with a few pixels altered, called sparse fooling images (SFIs). We first prove that SFIs always exist under mild conditions for linear and nonlinear models and reveal that complex models are more likely to be vulnerable to SFI attacks. Using two SFI generation methods, we demonstrate that in deeper layers, SFIs have features similar to those of natural images. Therefore, they fool DNNs successfully. Among the other layers, we discover that the max-pooling layer causes vulnerability to SFIs. The defense against SFIs and transferability are also discussed. This study highlights a new vulnerability of DNNs by introducing a novel class of images that are distributed extremely far from natural images. Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki |
Pattern Recognit. Lett. | 3 |
| 2023 | Crowd-Powered Photo Enhancement Featuring an Active Learning Based Local FilterabstractIn this study, we address local photo enhancement to improve the aesthetic quality of an input image by applying different effects to different regions. Existing photo enhancement methods are either not content-aware or not local; therefore, we propose a crowd-powered local enhancement method for content-aware local enhancement, which is achieved by asking crowd workers to locally optimize parameters for image editing functions. To make it easier to locally optimize the parameters, we propose an active learning based local filter. The parameters need to be determined at only a few key pixels selected by an active learning method, and the parameters at the other pixels are automatically predicted using a regression model. The parameters at the selected key pixels are independently optimized, breaking down the optimization problem into a sequence of single-slider adjustments. Our experiments show that the proposed filter outperforms existing filters, and our enhanced results are more visually pleasing than the results by the existing enhancement methods. Our source code and results are available athttps://github.com/satoshi-kosugi/crowd-powered. Satoshi Kosugi, Toshihiko Yamasaki |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Toward Extremely Lightweight Distracted Driver Recognition With Distillation-Based Neural Architecture Search and Knowledge TransferabstractThe number of traffic accidents has been continuously increasing in recent years worldwide. Many accidents are caused by distracted drivers, who take their attention away from driving. Motivated by the success of Convolutional Neural Networks (CNNs) in computer vision, many researchers developed CNN-based algorithms to recognize distracted driving from a dashcam and warn the driver against unsafe behaviors. However, current models have too many parameters, which is unfeasible for vehicle-mounted computing. This work proposes a novel knowledge-distillation-based framework to solve this problem. The proposed framework first constructs a high-performance teacher network by progressively strengthening the robustness to illumination changes from shallow to deep layers of a CNN. Then, the teacher network is used to guide the architecture searching process of a student network through knowledge distillation. After that, we use the teacher network again to transfer knowledge to the student network by knowledge distillation. Experimental results on the Statefarm Distracted Driver Detection Dataset and AUC Distracted Driver Dataset show that the proposed approach is highly effective for recognizing distracted driving behaviors from photos: (i) the teacher network’s accuracy surpasses the previous best accuracy; (ii) the student network achieves very high accuracy with only 0.42M parameters (around 55% of the previous most lightweight model). Furthermore, the student network architecture can be extended to a spatial-temporal 3D CNN for recognizing distracted driving from video clips. The 3D student network largely surpasses the previous best accuracy with only 2.03M parameters on the Drive&Act Dataset. The source code is available athttps://github.com/Dichao-Liu/Lightweight_Distracted_Driver_Recognition_with_Distillation-Based_NAS_and_Knowledge_Transfer Dichao Liu, Toshihiko Yamasaki, Yu Wang 0018, Kenji Mase, Jien Kato |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Subjective Functionality and Comfort Prediction for Apartment Floor Plans and Its Application to Intuitive Online Property SearchabstractThis paper presents a new user experience for online apartment search using functionality and comfort as query items. Specifically, it has three technical contributions. First, we present a new dataset on the perceived functionality and comfort scores of residential floor plans using nine question statements about the level of comfort, openness, privacy, etc. Second, we propose an algorithm to predict the scores from the floor plan images. Lastly, we implement a new apartment search system and conduct a large-scale usability study using crowdsourcing. The experimental results show that our apartment search system can provide a better user experience. To the best of our knowledge, this is the first work to propose a highly accurate machine learning model for predicting the subjective functionality and comfort of apartments. Taro Narahara, Toshihiko Yamasaki |
IEEE Trans. Multim. | 2 |
| 2022 | Fine-Grained Image Style Transfer with Visual Transformers
Huan Yang 0005, Jianlong Fu, Toshihiko Yamasaki, Baining Guo |
ACCV (3) | 4 |
| 2022 | Learning Where to Learn in Cross-View Self-Supervised LearningabstractSelf-supervised learning (SSL) has made enormous progress and largely narrowed the gap with the supervised ones, where the representation learning is mainly guided by a projection into an embedding space. During the projection, current methods simply adopt uniform aggregation of pixels for embedding; however, this risks involving object-irrelevant nuisances and spatial misalignment for different augmentations. In this paper, we present a new approach, Learning Where to Learn (LEWEL), to adaptively aggregate spatial information of features, so that the projected embeddings could be exactly aligned and thus guide the feature learning better. Concretely, we reinterpret the projection head in SSL as a per-pixel projection and predict a set of spatial alignment maps from the original features by this weight-sharing projection head. A spectrum of aligned embeddings is thus obtained by aggregating the features with spatial weighting according to these alignment maps. As a result of this adaptive alignment, we observe substantial improvements on both image-level prediction and dense prediction at the same time: LEWEL improves MoCov2 [15] by 1.6%/1.3%/0.5%/0.4% points, improves BYOL [14] by 1.3%/1.3%/0.7%/0.6% points, on ImageNet linear/semi-supervised classification, Pascal VOC semantic segmentation, and object detection, respectively.††Code: https://t.1y/ZI0A. Lang Huang 0001, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Toshihiko Yamasaki |
CVPR | 6 |
| 2022 | Detecting Deepfakes with Self-Blended ImagesabstractIn this paper, we present novel synthetic training data called self-blended images (SBIs) to detect deepfakes. SBIs are generated by blending pseudo source and target images from single pristine images, reproducing common forgery artifacts (e.g., blending boundaries and statistical inconsistencies between source and target images). The key idea behind SBIs is that more general and hardly recognizable fake samples encourage classifiers to learn generic and robust representations without overfitting to manipulation-specific artifacts. We compare our approach with state-of-the-art methods on FF++, CDF, DFD, DFDC, DFDCP, and FFIW datasets by following the standard cross-dataset and cross-manipulation protocols. Extensive experiments show that our method improves the model generalization to unknown manipulations and scenes. In particular, on DFDC and DFDCP where existing methods suffer from the domain gap between the training and test sets, our approach outperforms the baseline by 4.90% and 11.78% points in the cross-dataset evaluation, respectively. Code is available at https://github.com/mapooon/SelfBlendedImages. Kaede Shiohara, Toshihiko Yamasaki |
CVPR | 2 |
| 2022 | Improving Robustness to out-of-Distribution Data by Frequency-Based AugmentationabstractAlthough Convolutional Neural Networks (CNNs) have high accuracy in image recognition, they are vulnerable to adversarial examples and out-of-distribution data, and the difference from human recognition has been pointed out. In order to improve the robustness against out-of-distribution data, we present a frequency-based data augmentation technique that replaces the frequency components with other images of the same class. When the training data are CIFAR10 and the out-of-distribution data are SVHN, the Area Under Receiver Operating Characteristic (AUROC) curve of the model trained with the proposed method increases from 89.22% to 98.15%, and further increased to 98.59% when combined with another data augmentation method. Furthermore, we experimentally demonstrate that the robust model for out-of-distribution data uses a lot of high-frequency components of the image. Koki Mukai, Soichiro Kumano, Toshihiko Yamasaki |
ICIP | 3 |
| 2022 | Sat: Self-Adaptive Training for Fashion Compatibility PredictionabstractThis paper presents a self-adaptive training (SAT) model for fashion compatibility prediction. It focuses on the learning of some hard items, such as those that share similar color, texture, and pattern features but are considered incompatible due to the aesthetics or temporal shifts. Specifically, we first design a method to define hard outfits and a difficulty score (DS) is defined and assigned to each outfit based on the difficulty in recommending an item for it. Then, we propose a self-adaptive triplet loss (SATL), where the DS of the outfit is considered. Finally, we propose a very simple conditional similarity network combining the proposed SATL to achieve the learning of hard items in the fashion compatibility prediction. Experiments on the publicly available Polyvore Outfits and Polyvore Outfits-D datasets demonstrate our SAT’s effectiveness in fashion compatibility prediction. Besides, our SATL can be easily extended to other conditional similarity networks to improve their performance. Ling Xiao 0001, Toshihiko Yamasaki |
ICIP | 2 |
| 2022 | Graph Neural Network Based Living Comfort Prediction Using Real Estate Floor Plan ImagesabstractIn recent years, machine learning has been widely used in the real estate field. However, most of these previous studies have been limited to analysis based on objective perspectives, such as analysis of the structure of the floor plan and rent estimation. On the other hand, we focus on the subjective "living comfort" of real estate properties and aim to predict people's impressions of properties based on information obtained from floor plan images. Specifically, by using deep learning to analyze floor plan images and graph structures reflecting the floor plans, it becomes possible to predict the attractiveness of each property in terms of spaciousness, modernity, privacy, and so on. As a result of the experiments, the effectiveness of using both the floor plan image and the corresponding graph structure for prediction was confirmed. Ryota Kitabayashi, Taro Narahara, Toshihiko Yamasaki |
MMAsia | 3 |
| 2022 | Green Hierarchical Vision Transformer for Masked Image ModelingabstractWe present an efficient approach for Masked Image Modeling (MIM) with hierarchical Vision Transformers (ViTs), allowing the hierarchical ViTs to discard masked patches and operate only on the visible ones. Our approach consists of three key designs. First, for window attention, we propose a Group Window Attention scheme following the Divide-and-Conquer strategy. To mitigate the quadratic complexity of the self-attention w.r.t. the number of patches, group attention encourages a uniform partition that visible patches within each local window of arbitrary size can be grouped with equal size, where masked self-attention is then performed within each group. Second, we further improve the grouping strategy via the Dynamic Programming algorithm to minimize the overall computation cost of the attention on the grouped patches. Third, as for the convolution layers, we convert them to the Sparse Convolution that works seamlessly with the sparse data, i.e., the visible patches in MIM. As a result, MIM can now work on most, if not all, hierarchical ViTs in a green and efficient way. For example, we can train the hierarchical ViTs, e.g., Swin Transformer and Twins Transformer, about 2.7$\times$ faster and reduce the GPU memory usage by 70%, while still enjoying competitive performance on ImageNet classification and the superiority on downstream COCO object detection benchmarks. Lang Huang 0001, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Toshihiko Yamasaki |
NeurIPS | 6 |
| 2022 | Cover: International Journal of Intelligent Systems, Volume 37 Issue 5 May 2022abstractCover Caption: The cover image is based on the Research Article Efficient virtual data search for annotationfree vehicle reidentification by Zhijing Wan et al., https://doi.org/10.1002/int.22829. Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu |
Int. J. Intell. Syst. | 4 |
| 2022 | Efficient virtual data search for annotation-free vehicle reidentificationabstractVehicle reidentification (re-ID) is the task of retrieving the same vehicle across nonoverlapping cameras, which has made significant progress with the help of abundant manually annotated real images. To avoid the time-consuming and tedious labeling of real images, virtual data sets with large-scale synthetic images have recently been constructed to perform annotation-free model training. However, current methods fail to exploit the potential of virtual data search, that is, searching valuable and representative virtual subdata set for efficient training. This paper presents a novel data sampling strategy from both semantic and feature levels to perform an effective data search. The semantic level determines the sample number of each vehicle identity via the consistency constraint of attribute distribution for source domain and target domain; while the feature level searches valuable and representative samples of each vehicle identity. To our knowledge, we are among the first attempts to search effective virtual data to perform annotation-free vehicle re-ID. Extensive cross-domain experiments from virtual vehicle re-ID data sets to real vehicle re-ID data sets show that our data sampling strategy can significantly reduce the training data volume and even boost the re-ID performance. Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu |
Int. J. Intell. Syst. | 4 |
| 2022 | Spatially adaptive multi-scale contextual attention for image inpaintingabstractAbstract Image inpainting is the task to fill missing regions of an image. Recently, researchers have achieved a great performance by using convolutional neural networks (CNNs) with the conventional patch-matching method. Existing methods compute the attention scores, which are based on the similarity of patches between the known and missing regions. Considering that patches at different spatial positions can convey different levels of detail, we propose a spatially adaptive multi-scale attention score that uses the patches of different scales to compute scores for each pixel at different positions. Through experiments on the Paris Street View and Places datasets, our proposal shows slight improvement compared with some related methods on the quantitative evaluation metrics commonly used in the existing methods. Moreover, we found that these quantitative metrics are not appropriate enough considering the subjective impressions of the generated images. Therefore, we conducted subjective evaluation through user study for comparison, which shows that our proposal has superiority of performance generating much more detailed and subjectively plausible images. Toshihiko Yamasaki |
Multim. Tools Appl. | 3 |
| 2022 | An Improved Inter-Intra Contrastive Learning Framework on Self-Supervised Video RepresentationabstractIn this paper, we propose a self-supervised contrastive learning method to learn video feature representations. In traditional self-supervised contrastive learning methods, constraints from anchor, positive, and negative data pairs are used to train the model. In such a case, different samplings of the same video are treated as positives, and video clips from different videos are treated as negatives. Because the spatio-temporal information is important for video representation, we set the temporal constraints more strictly by introducing intra-negative samples. In addition to samples from different videos, negative samples are extended by breaking temporal relations in video clips from the same anchor video. With the proposed Inter-Intra Contrastive (IIC) framework, we can train spatio-temporal convolutional networks to learn feature representations from videos. Strong data augmentations, residual clips, as well as head projector are utilized to construct an improved version. Three kinds of intra-negative generation functions are proposed and extensive experiments using different network backbones are conducted on benchmark datasets. Without using pre-computed optical flow data, our improved version can outperform previous IIC by a large margin, such as 19.4% (from 36.8% to 56.2%) and 5.2% (from 15.5% to 20.7%) points improvements in top-1 accuracy on UCF101 and HMDB51 datasets for video retrieval, respectively. For video recognition, over 3% points improvements can also be obtained on these two benchmark datasets. Discussions and visualizations validate that our IICv2 can capture better temporal clues and indicate the potential mechanism. Toshihiko Yamasaki |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Sampling and Re-Weighting: Towards Diverse Frame Aware Unsupervised Video Person Re-IdentificationabstractVideo person re-identification (re-ID) methods extract richer features from video tracklets than image-based ones and have received growing attention. However, existing supervised methods require numerous cross-camera identity labels, which is impractical for large-scale data. Although clustering-based unsupervised methods have been exploited to obtain pseudo labels and train the models iteratively for video person re-ID, they remain in their infancy due to the diversity of person images and uncertainty in the image quality of video tracklets. In this work, we employ two strategies ofSampling andRe-weighting forClustering (SRC) to obtain robust and discriminative person feature representations. This method considers the influence of two kinds of frames in the tracklet: 1) Detection errors or heavy occlusions generate noisy frames in the tracklet. These tracklets with noisy frames may be assigned with unreliable annotations during clustering. 2) Different frames are identified by the model with varying degrees of difficulty, caused by pose changes or partial occlusions. We call them hard frames, which are hard to identify but informative. To alleviate these problems, we propose a dynamic noise trimming module and diverse frame re-weighting module for sampling and re-weighting. The dynamic noise trimming module strengthens the dependability of the tracklet representation by removing noisy frames to enhance the clustering accuracy. The diverse frame re-weighting module focuses on training hard frames to enhance the learning of rich information from tracklet. Experiments on three video datasets,i.e.DukeMTMC-VideoReID, MARS and PRID2011, demonstrate the effectiveness of the proposed SRC under the unsupervised re-ID setting. Pengyu Xie, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki |
IEEE Trans. Multim. | 4 |
| 2021 | Very Important Person Localization in Unconstrained Conditions: A New BenchmarkabstractThis paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, current datasets: 1) are limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relatively simple, and the face of VIP is always in frontal view and salient. To tackle these problems, the proposed Unconstrained-7k dataset is featured in two aspects. First, it contains over 7,000 annotated images, making it the largest VIPLoc dataset under unconstrained conditions to date. Second, our dataset is collected freely on the Internet, including multiple scenes, where images are in unconstrained conditions. VIPs in the new dataset are in different settings, e.g., large view variation, varying sizes, occluded, and complex scenes. Meanwhile, each image has more persons (> 20), making the dataset more challenging. As a minor contribution, motivated by the observation that VIPs are highly related to not only neighbors but also iconic objects, this paper proposes a Joint Social Relation and Individual Interaction Graph Neural Networks (JSRII-GNN) for VIPLoc. Experiments show that the JSRII-GNN yields competitive accuracy on NCAA (National Collegiate Athletic Association), MS (Multi-scene), and Unconstrained-7k datasets. https://github.com/xiaowang1516/VIPLoc. Xiao Wang 0029, Zheng Wang 0007, Toshihiko Yamasaki, Wenjun Zeng 0001 |
AAAI | 3 |
| 2021 | Unsupervised Video Person Re-Identification via Noise and Hard Frame Aware ClusteringabstractUnsupervised video-based person re-identification (re-ID) methods extract richer features from video tracklets than image-based ones. The state-of-the-art methods utilize clustering to obtain pseudo-labels and train the models iteratively. However, they underestimate the influence of two kinds of frames in the tracklet: 1) noise frames caused by detection errors or heavy occlusions exist in the tracklet, which may be allocated with unreliable labels during clustering; 2) the tracklet also contains hard frames caused by pose changes or partial occlusions, which are difficult to distinguish but informative. This paper proposes a Noise and Hard frame Aware Clustering (NHAC) method. NHAC consists of a graph trimming module and a node re-sampling module. The graph trimming module obtains stable graphs by removing noise frame nodes to improve the clustering accuracy. The node re-sampling module enhances the training of hard frame nodes to learn rich tracklet information. Experiments conducted on two video-based datasets demonstrate the effectiveness of the proposed NHAC under the unsupervised re-ID setting. Pengyu Xie, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki |
ICME | 4 |
| 2021 | Location Predicts You: Location Prediction via Bi-direction Speculation and Dual-level AssociationabstractLocation prediction is of great importance in location-based applications for the construction of the smart city. To our knowledge, existing models for location prediction focus on the users' preference on POIs from the perspective of the human side. However, modeling users' interests from the historical trajectory is still limited by the data sparsity. Additionally, most of existing methods predict the next location according to the individual data independently. But the data sparsity makes it difficult to mine explicit mobility patterns or capture the casual behavior for each user. To address the issues above, we propose a novel Bi-direction Speculation and Dual-level Association method (BSDA), which considers both users' interests in POIs and POIs' appeal to users. Furthermore, we develop the cross-user and cross-POI association to alleviate the data sparsity by similar users and POIs to enrich the candidates. Experimental results on two public datasets demonstrate that BSDA achieves significant improvements over state-of-the-art methods. Ruimin Hu, Zheng Wang 0007, Toshihiko Yamasaki |
IJCAI | 4 |
| 2021 | Edge-Level Explanations for Graph Neural Networks by Extending Explainability Methods for Convolutional Neural NetworksabstractGraph Neural Networks (GNNs) are deep learning models that take graph data as inputs, and they are applied to various tasks such as traffic prediction and molecular property prediction. However, owing to the complexity of the GNNs, it has been difficult to analyze which parts of inputs affect the GNN model’s outputs. In this study, we extend explainability methods for Convolutional Neural Networks (CNNs), such as Local Interpretable Model-Agnostic Explanations (LIME), Gradient-Based Saliency Maps, and Gradient-Weighted Class Activation Mapping (Grad-CAM) to GNNs, and predict which edges in the input graphs are important for GNN decisions. The experimental results indicate that the LIME-based approach is the most efficient explainability method for multiple tasks in the real-world situation, outperforming even the state-of-the-art method in GNN explainability. Tetsu Kasanishi, Toshihiko Yamasaki |
ISM | 3 |
| 2021 | Sequential Banner Design Optimization with Deep Reinforcement LearningabstractMany banner images in web advertising are composed of multiple image elements. In this study, we propose a method to optimize the placement of each image element sequentially by self-supervised reinforcement learning. The environment randomly disturbs the position of image elements at first; then, the agent learns how to correct it in a self-supervision manner, enabling the agent to learn how to place each image element for good banner design. For smooth training, we designed new rewards based on the absolute and relative position of image elements. In addition, by incorporating the output of the click-through rate (CTR) prediction model into the reward, the optimized banner design can be expected to have a higher CTR. Yusuke Kondo, Hiroyuki Seshime, Toshihiko Yamasaki |
ISM | 4 |
| 2021 | Reproducibility Companion Paper: Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkabstractIn this companion paper, we provide details of the artifacts to support the replication of "Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework", which was presented at MM'20. The Inter-intra Contrastive (IIC) framework aims to extract more discriminative temporal information by extending intra-negative samples in contrastive self-supervised learning. In this paper, we first summarize our contribution. Then we explain the file structure of the source code and detailed settings. Since our proposal is a framework which contain a lot of different settings, we provide some custom settings to help other researchers to use our methods easily. The source code is available at https://github.com/BestJuly/IIC. Toshihiko Yamasaki, Jingjing Chen 0001, Steven Alexander Hicks |
ACM Multimedia | 3 |
| 2021 | Style-Aware Image Recommendation for Social Media MarketingabstractSocial media have become a popular platform for brands to allocate marketing budget and build their relationship with customers. Posting images with a consistent concept on social media helps customers recognize, remember, and consider brands. This strategy is known as brand concept consistency in marketing literature. Consequently, brands spend immense manpower and financial resources in choosing which images to post or repost. Therefore, automatically recommending images with a consistent brand concept is a necessary task for social media marketing. In this paper, we propose a content-based recommendation system that learns the concept of brands and recommends images that are coherent with the brand. Specifically, brand representation is performed from the brand posts on social media. Existing methods rely on visual features extracted by pre-trained neural networks, which can represent objects in the image but not the style of the image. To bridge this gap, a framework using both object and style vectors as input is proposed to learn the brand representation. In addition, we show that the proposed method can not only be applied to brands but also be applied to influencers. We collected a new Instagram influencer dataset, consisting of 616 influencers and about 1 million images, which can greatly benefit future research in this area. The experimental results on two large-scale Instagram datasets show the superiority of the proposed method over state-of-the-art methods. Yiwei Zhang 0014, Toshihiko Yamasaki |
ACM Multimedia | 2 |
| 2021 | Learning From Synthetic Shadows for Shadow Detection and RemovalabstractShadow removal is an essential task in computer vision and computer graphics. Recent shadow removal approaches all train convolutional neural networks (CNN) on real paired shadow/shadow-free or shadow/shadow-free/mask image datasets. However, obtaining a large-scale, diverse, and accurate dataset has been a big challenge, and it limits the performance of the learned models on shadow images with unseen shapes/intensities. To overcome this challenge, we present SynShadow, a novel large-scale synthetic shadow/shadow-free/matte image triplets dataset and a pipeline to synthesize it. We extend a physically-grounded shadow illumination model and synthesize a shadow image given an arbitrary combination of a shadow-free image, a matte image, and shadow attenuation parameters. Owing to the diversity, quantity, and quality of SynShadow, we demonstrate that shadow removal models trained on SynShadow perform well in removing shadows with diverse shapes and intensities on some challenging benchmarks. Furthermore, we show that merely fine-tuning from a SynShadow-pre-trained model improves existing shadow detection and removal models. Codes are publicly available athttps://github.com/naoto0804/SynShadow. Naoto Inoue, Toshihiko Yamasaki |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Rethinking Motion Representation: Residual Frames With 3D ConvNetsabstractRecently, 3D convolutional networks yield good performance in action recognition. However, an optical flow stream is still needed for motion representation to ensure better performance, whose cost is very high. In this paper, we propose a cheap but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 35.6% and 26.6% points improvements over top-1 accuracy can be achieved on the UCF101 and HMDB51 datasets when trained from scratch using ResNet-18-3D. We deeply analyze the effectiveness of this modality compared to normal RGB video clips, and find that better motion features can be extracted using residual frames with 3D ConvNets. Considering that residual frames contain little information of object appearance, we further use a 2D convolutional network to extract appearance features and combine them together to form a two-path solution. In this way, we can achieve better performance than some methods which even used an additional optical flow stream. Moreover, the proposed residual-input path can outperform RGB counterpart on unseen datasets when we apply trained models to video retrieval tasks. Huge improvements can also be obtained when the residual inputs are applied to video-based self-supervised learning methods, revealing better motion representation and generalization ability of our proposal. Toshihiko Yamasaki |
IEEE Trans. Image Process. | 3 |
| 2021 | Semantic Explanation for Deep Neural Networks Using Feature InteractionsabstractGiven the promising results obtained by deep-learning techniques in multimedia analysis, the explainability of predictions made by networks has become important in practical applications. We present a method to generate semantic and quantitative explanations that are easily interpretable by humans. The previous work to obtain such explanations has focused on the contributions of each feature, taking their sum to be the prediction result for a target variable; the lack of discriminative power due to this simple additive formulation led to low explanatory performance. Our method considers not only individual features but also their interactions, for a more detailed interpretation of the decisions made by networks. The algorithm is based on the factorization machine, a prediction method that calculates factor vectors for each feature. We conducted experiments on multiple datasets with different models to validate our method, achieving higher performance than the previous work. We show that including interactions not only generates explanations but also makes them richer and is able to convey more information. We show examples of produced explanations in a simple visual format and verify that they are easily interpretable and plausible. Bohui Xia, Toshihiko Yamasaki |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Unpaired Image Enhancement Featuring Reinforcement-Learning-Controlled Image Editing SoftwareabstractThis paper tackles unpaired image enhancement, a task of learning a mapping function which transforms input images into enhanced images in the absence of input-output image pairs. Our method is based on generative adversarial networks (GANs), but instead of simply generating images with a neural network, we enhance images utilizing image editing software such as Adobe® Photoshop® for the following three benefits: enhanced images have no artifacts, the same enhancement can be applied to larger images, and the enhancement is interpretable. To incorporate image editing software into a GAN, we propose a reinforcement learning framework where the generator works as the agent that selects the software's parameters and is rewarded when it fools the discriminator. Our framework can use high-quality non-differentiable filters present in image editing software, which enables image enhancement with high performance. We apply the proposed method to two unpaired image enhancement tasks: photo enhancement and face beautification. Our experimental results demonstrate that the proposed method achieves better performance, compared to the performances of the state-of-the-art methods based on unpaired learning. Satoshi Kosugi, Toshihiko Yamasaki |
AAAI | 2 |
| 2020 | PresentationTrainer: Oral Presentation Support System for Impression-Related FeedbackabstractIn order to support the pratice of oral presentation, we developed PresentationTrainer which includes (1) a presentation impression prediction system and (2) a presentation slide analysis system. For the presentation impression prediction system, we proposed two methods, using Support Vector Machine and Markov Random Field, or using multimodal neural network, to predict audiences' impressions for speech videos. For the slide analysis system, we used Convolutional Neural Network and Global Average Pooling to evaluate the design of slides. We then used Class Activation Mapping to provide visual feedback for showing which areas should be modified. Shengzhou Yi, Hiroshi Yumoto, Toshihiko Yamasaki |
AAAI | 4 |
| 2020 | Investigating Generalization in Neural Networks Under Optimally Evolved Training PerturbationsabstractIn this paper, we study the generalization properties of neural networks under input perturbations and show that minimal training data corruption by a few pixel modifications can cause drastic overfitting. We propose an evolutionary algorithm to search for optimal pixel perturbations using novel cost function inspired from literature in domain adaptation that explicitly maximizes the generalization gap and domain divergence between clean and corrupted images. Our method outperforms previous pixel-based data distribution shift methods on state-of-the-art Convolutional Neural Networks (CNNs) architectures. Interestingly, we find that the choice of optimization plays an important role in generalization robustness due to the empirical observation that SGD is resilient to such training data corruption unlike adaptive optimization techniques (ADAM). Subhajit Chaudhury, Toshihiko Yamasaki |
ICASSP | 2 |
| 2020 | Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown ObjectsabstractHand pose estimation is important for human computer interaction, but the performance is not satisfying when the hand is interacting with objects. To alleviate the influence of unknown objects, we propose a novel weakly supervised segmentation guided scheme to estimate hand poses. Approximate hand masks generated from annotations of sparse hand joints are used to supervise the segmentation task. Better features can be extracted since they are shared between the two tasks of hand segmentation and hand pose estimation. With the guidance of weakly supervised segmentation, the network can learn intermediate features balanced between focusing on the foreground and preserving contextual information. Finally the xy and z coordinates are estimated in different branches but utilizing shared feature maps. Experimental results of three different tasks on the publicly available FHAD dataset demonstrate the effectiveness of the proposed architecture. Cairong Zhang, Guijin Wang, Xinghao Chen 0001, Pengwei Xie, Toshihiko Yamasaki |
ICASSP | 5 |
| 2020 | Motion Representation Using Residual Frames with 3D CNNabstractRecently, 3D convolutional networks (3D ConvNets) yield good performance in action recognition. However, optical flow stream is still needed to ensure better performance, the cost of which is very high. In this paper, we propose a fast but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 35.6% and 26.6% points improvements over top-l accuracy can be obtained on the UCF101 and HMDB51 datasets when ResNet-18 models are trained from scratch. And we achieved the state-of-the-art results in this training mode. Analysis shows that better motion features can be extracted using residual frames compared to RGB counterpart. By combining with a simple appearance path, our proposal can be even better than some methods using optical flow streams. Toshihiko Yamasaki |
ICIP | 3 |
| 2020 | Feature Point Matching in Cross-Spectral Images with Cycle Consistency LearningabstractFeature point matching is an important problem because its applications cover a wide range of tasks in computer vision. Deep learning-based methods for learning local features have recently shown superior performance. However, it is not easy to collect the training data in these methods, especially in cross-spectral settings such as the correspondence between RGB and near-infrared images. In this paper, we propose an unsupervised learning method for general feature point matching. Because we train a convolutional neural network as a feature extractor in order to satisfy the cycle consistency of the correspondences between an input image pair, the proposed method does not require supervision and works even in cross-spectral settings. In our experiments, we apply the proposed method to stereo matching, which is a dense feature point matching problem. The experimental results, which simulate cross-spectral settings with three different settings, i.e., RGB stereo, RGB vs gray-scale, and anaglyph (red vs cyan), show that our proposed method outperforms the compared methods, which employ handcrafted features for stereo matching, by a significant margin. Ryosuke Furuta, Naoaki Noguchi, Toshihiko Yamasaki |
ICPR | 4 |
| 2020 | Predicting Online Video Advertising Effects with Multimodal Deep LearningabstractWith expansion of the video advertising market, research to predict the effects of video advertising is getting more attention. Although effect prediction of image advertising has been explored a lot, prediction for video advertising is still challenging with seldom research. In this research, we propose a method for predicting the click through rate (CTR) of video advertisements and analyzing the factors that determine the CTR. In this paper, we demonstrate an optimized framework for accurately predicting the effects by taking advantage of the multimodal nature of online video advertisements including video, text, and metadata features. In particular, the two types of metadata, i.e., categorical and continuous, are properly separated and normalized. To avoid overfitting, which is crucial in our task because the training data are not very rich, additional regularization layers are inserted. Experimental results show that our approach can achieve a correlation coefficient as high as 0.695, which is a significant improvement from the baseline (0.487). Jun Ikeda, Hiroyuki Seshime, Toshihiko Yamasaki |
ICPR | 4 |
| 2020 | MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020abstractThe International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173 Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki |
ICMR | 5 |
| 2020 | Self-Play Reinforcement Learning for Fast Image RetargetingabstractIn this study, we address image retargeting, which is a task that adjusts input images to arbitrary sizes. In one of the best-performing methods called MULTIOP, multiple retargeting operators were combined and retargeted images at each stage were generated to find the optimal sequence of operators that minimized the distance between original and retargeted images. The limitation of this method is in its tremendous processing time, which severely prohibits its practical use. Therefore, the purpose of this study is to find the optimal combination of operators within a reasonable processing time; we propose a method of predicting the optimal operator for each step using a reinforcement learning agent. The technical contributions of this study are as follows. Firstly, we propose a reward based on self-play, which will be insensitive to the large variance in the content-dependent distance measured in MULTIOP. Secondly, we propose to dynamically change the loss weight for each action to prevent the algorithm from falling into a local optimum and from choosing only the most frequently used operator in its training. Our experiments showed that we achieved multi-operator image retargeting with less processing time by three orders of magnitude and the same quality as the original multi-operator-based method, which was the best-performing algorithm in retargeting tasks. Nobukatsu Kajiura, Satoshi Kosugi, Toshihiko Yamasaki |
ACM Multimedia | 4 |
| 2020 | Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkabstractWe propose a self-supervised method to learn feature representations from videos. A standard approach in traditional self-supervised methods uses positive-negative data pairs to train with contrastive learning strategy. In such a case, different modalities of the same video are treated as positives and video clips from a different video are treated as negatives. Because the spatio-temporal information is important for video representation, we extend the negative samples by introducing intra-negative samples, which are transformed from the same anchor video by breaking temporal relations in video clips. With the proposed Inter-Intra Contrastive (IIC) framework, we can train spatio-temporal convolutional networks to learn video representations. There are many flexible options in our IIC framework and we conduct experiments by using several different configurations. Evaluations are conducted on video retrieval and video recognition tasks using the learned video representation. Our proposed IIC outperforms current state-of-the-art results by a large margin, such as 16.7% and 9.5% points improvements in top-1 accuracy on UCF101 and HMDB51 datasets for video retrieval, respectively. For video recognition, improvements can also be obtained on these two benchmark datasets. Toshihiko Yamasaki |
ACM Multimedia | 3 |
| 2020 | RGB2AO: Ambient Occlusion Generation from RGB ImagesabstractAbstract We present RGB2AO, a novel task to generate ambient occlusion (AO) from a single RGB image instead of screen space buffers such as depth and normal. RGB2AO produces a new image filter that creates a non‐directional shading effect that darkens enclosed and sheltered areas. RGB2AO aims to enhance two 2D image editing applications: image composition and geometry‐aware contrast enhancement. We first collect a synthetic dataset consisting of pairs of RGB images and AO maps. Subsequently, we propose a model for RGB2AO by supervised learning of a convolutional neural network (CNN), considering 3D geometry of the input image. Experimental results quantitatively and qualitatively demonstrate the effectiveness of our model. Naoto Inoue, Daichi Ito, Yannick Hold-Geoffroy, Long Mai, Brian L. Price, Toshihiko Yamasaki |
Comput. Graph. Forum | 6 |
| 2020 | PixelRL: Fully Convolutional Network With Reinforcement Learning for Image ProcessingabstractThis article tackles a new problem setting: reinforcement learning with pixel-wise rewards (pixelRL) for image processing. After the introduction of the deep Q-network, deep RL has been achieving great success. However, the applications of deep reinforcement learning (RL) for image processing are still limited. Therefore, we extend deep RL to pixelRL for various image processing applications. In pixelRL, each pixel has an agent, and the agent changes the pixel value by taking an action. We also propose an effective learning method for pixelRL that significantly improves the performance by considering not only the future states of the own pixel but also those of the neighbor pixels. The proposed method can be applied to some image processing tasks that require pixel-wise manipulations, where deep RL has never been applied. Besides, it is possible to visualize what kind of operation is employed for each pixel at each iteration, which would help us understand why and how such an operation is chosen. We also believe that our technology can enhance the explainability and interpretability of the deep neural networks. In addition, because the operations executed at each pixels are visualized, we can change or modify the operations if necessary. We apply the proposed method to a variety of image processing tasks: image denoising, image restoration, local color enhancement, and saliency-driven image editing. Our experimental results demonstrate that the proposed method achieves comparable or better performance, compared with the state-of-the-art methods based on supervised learning. The source code is available on https://github.com/rfuruta/pixelRL. Ryosuke Furuta, Naoto Inoue, Toshihiko Yamasaki |
IEEE Trans. Multim. | 3 |
| 2019 | Fully Convolutional Network with Multi-Step Reinforcement Learning for Image ProcessingabstractThis paper tackles a new problem setting: reinforcement learning with pixel-wise rewards (pixelRL) for image processing. After the introduction of the deep Q-network, deep RL has been achieving great success. However, the applications of deep RL for image processing are still limited. Therefore, we extend deep RL to pixelRL for various image processing applications. In pixelRL, each pixel has an agent, and the agent changes the pixel value by taking an action. We also propose an effective learning method for pixelRL that significantly improves the performance by considering not only the future states of the own pixel but also those of the neighbor pixels. The proposed method can be applied to some image processing tasks that require pixel-wise manipulations, where deep RL has never been applied.We apply the proposed method to three image processing tasks: image denoising, image restoration, and local color enhancement. Our experimental results demonstrate that the proposed method achieves comparable or better performance, compared with the state-of-the-art methods based on supervised learning. Ryosuke Furuta, Naoto Inoue, Toshihiko Yamasaki |
AAAI | 3 |
| 2019 | Object-Aware Instance Labeling for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD), where a detector is trained with only image-level annotations, is attracting more and more attention. As a method to obtain a well-performing detector, the detector and the instance labels are updated iteratively. In this study, for more efficient iterative updating, we focus on the instance labeling problem, a problem of which label should be annotated to each region based on the last localization result. Instead of simply labeling the top-scoring region and its highly overlapping regions as positive and others as negative, we propose more effective instance labeling methods as follows. First, to solve the problem that regions covering only some parts of the object tend to be labeled as positive, we find regions covering the whole object focusing on the context classification loss. Second, considering the situation where the other objects contained in the image can be labeled as negative, we impose a spatial restriction on regions labeled as negative. Using these instance labeling methods, we train the detector on the PASCAL VOC 2007 and 2012 and obtain significantly improved results compared with other state-of-the-art approaches. Satoshi Kosugi, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICCV | 2 |
| 2019 | Reinforcing the Robustness of a Deep Neural Network to Adversarial Examples by Using Color Quantization of Training Image DataabstractRecent works have shown the vulnerability of deep convolutional neural network (DCNN) to adversarial examples with malicious perturbations. In particular, Black-Box attacks without information of parameter and architectures of the target models are feared as realistic threats. To address this problem, we propose a method using an ensemble of models trained by color-quantized data with loss maximization. Color-quantization can allow the trained models to focus on learning conspicuous spatial features to enhance the robustness of DCNNs to adversarial examples. The proposed method can be adapted to Black-Box attacks with no need of particular attack algorithm for the defense. The results of our experiments validated the effectiveness for preventing decrease in the test accuracy with adversarial perturbation. Shuntaro Miyazato, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 3 |
| 2019 | User-Aware Folk Popularity Rank: User-Popularity-Based Tag Recommendation That Can Enhance Social PopularityabstractIn this paper we propose a method that can enhance the social popularity of a post (i.e., the number of views or likes) by recommending appropriate hash tags considering both content popularity and user popularity. A previous approach called FolkPopularityRank (FP-Rank) considered only the relationship among images, tags, and their popularity. However, the popularity of an image/video is strongly affected by who uploaded it. Therefore, we develop an algorithm that can incorporate user popularity and users' tag usage tendency into the FP-Rank algorithm. The experimental results using 60,000 training images with their accompanying tags and 1,000 test data, which were actually uploaded to a real social network service (SNS), show that, in ten days, our proposed algorithm can achieve 1.2 times more views than the FP-Rank algorithm. This technology would be critical to individual users and companies/brands who want to promote themselves in SNSs. Yiwei Zhang 0014, Toshihiko Yamasaki |
ACM Multimedia | 3 |
| 2019 | Weakly Supervised Video Summarization by Hierarchical Reinforcement LearningabstractConventional video summarization approaches based on reinforcement learning have the problem that the reward can only be received after the whole summary is generated. Such kind of reward is sparse and it makes reinforcement learning hard to converge. Another problem is that labelling each shot is tedious and costly, which usually prohibits the construction of large-scale datasets. To solve these problems, we propose a weakly supervised hierarchical reinforcement learning framework, which decomposes the whole task into several subtasks to enhance the summarization quality. This framework consists of a manager network and a worker network. For each subtask, the manager is trained to set a subgoal only by a task-level binary label, which requires much fewer labels than conventional approaches. With the guide of the subgoal, the worker predicts the importance scores for video shots in the subtask by policy gradient according to both global reward and innovative defined sub-rewards to overcome the sparse problem. Experiments on two benchmark datasets show that our proposal has achieved the best performance, even better than supervised approaches. Toshihiko Yamasaki |
MMAsia | 4 |
| 2019 | Session details: Multimedia ServiceabstractNo abstract available. Toshihiko Yamasaki |
MMAsia | 1 |
| 2019 | Measuring Similarity between Brands using Followers' Post in Social MediaabstractIn this paper, we propose a new measure to estimate the similarity between brands via posts of brands' followers on social network services (SNS). Our method was developed with the intention of exploring the brands that customers are likely to jointly purchase. Nowadays, brands use social media for targeted advertising because influencing users' preferences can greatly affect the trends in sales. We assume that data on SNS allows us to make quantitative comparisons between brands. Our proposed algorithm analyzes the daily photos and hashtags posted by each brand's followers. By clustering them and converting them to histograms, we can calculate the similarity between brands. We evaluated our proposed algorithm with purchase logs, credit card information, and answers to the questionnaires. The experimental results show that the purchase data maintained by a mall or a credit card company can predict the co-purchase very well, but not the customer's willingness to buy products of new brands. On the other hand, our method can predict the users' interest on brands with a correlation value over 0.53, which is pretty high considering that such interest to brands are high subjective and individual dependent. Yiwei Zhang 0014, Yoshiaki Sakai, Toshihiko Yamasaki |
MMAsia | 4 |
| 2019 | Deep Feature Interaction Embedding for Pair Matching PredictionabstractOnline dating services have become popular in modern society. Pair matching prediction between two users in these services can help efficiently increase the possibility of finding their life partners. Deep learning based methods with automatic feature interaction functions such as Factorization Machines (FM) and cross network of Deep & Cross Network (DCN) can model sparse categorical features, which are effective to many recommendation tasks of web applications. To solve the partner recommendation task, we improve these FM-based deep models and DCN by enhancing the representation of feature interaction embedding and proposing a novel design of interaction layer avoiding information loss. Through the experiments on two real-world datasets of two online dating companies, we demonstrate the superior performances of our proposed designs. Luwei Zhang, Toshihiko Yamasaki |
MMAsia | 3 |
| 2019 | Learning to Trace: Expressive Line Drawing Generation from PhotographsabstractAbstract In this paper, we present a new computational method for automatically tracing high‐resolution photographs to create expressive line drawings. We define expressive lines as those that convey important edges, shape contours, and large‐scale texture lines that are necessary to accurately depict the overall structure of objects (similar to those found in technical drawings) while still being sparse and artistically pleasing. Given a photograph, our algorithm extracts expressive edges and creates a clean line drawing using a convolutional neural network (CNN). We employ an end‐to‐end trainable fully‐convolutional CNN to learn the model in a data‐driven manner. The model consists of two networks to cope with two sub‐tasks; extracting coarse lines and refining them to be more clean and expressive. To build a model that is optimal for each domain, we construct two new datasets for face/body and manga background. The experimental results qualitatively and quantitatively demonstrate the effectiveness of our model. We further illustrate two practical applications. Naoto Inoue, Daichi Ito, Ning Xu 0007, Brian L. Price, Toshihiko Yamasaki |
Comput. Graph. Forum | 6 |
| 2019 | Efficient and interactive spatial-semantic image retrievalabstractThis paper proposes an efficient image retrieval system. When users wish to retrieve images with semantic and spatial constraints ( e.g ., a horse is located at the center of the image, and a person is riding on the horse), it is difficult for conventional text-based retrieval systems to retrieve such images exactly. In contrast, the proposed system can consider both semantic and spatial information, because it is based on semantic segmentation using fully convolutional networks (FCN). The proposed system can accept three types of images as queries: a segmentation map sketched by the user, a natural image, or a combination of the two. The distance between the query and each image in the database is calculated based on the output probability maps from the FCN. In order to make the system efficient in terms of both the computational time and memory usage, we employ the product quantization (PQ) technique. The experimental results show that the PQ is compatible with the FCN-based image retrieval system, and that the quantization process results in little information loss. It is also shown that our method outperforms a conventional text-based search system. Ryosuke Furuta, Naoto Inoue, Toshihiko Yamasaki |
Multim. Tools Appl. | 3 |
| 2018 | Local and Global Optimization Techniques in Graph-Based ClusteringabstractThe goal of graph-based clustering is to divide a dataset into disjoint subsets with members similar to each other from an affinity (similarity) matrix between data. The most popular method of solving graph-based clustering is spectral clustering. However, spectral clustering has drawbacks. Spectral clustering can only be applied to macroaverage-based cost functions, which tend to generate undesirable small clusters. This study first introduces a novel cost function based on micro-average. We propose a local optimization method, which is widely applicable to graph-based clustering cost functions. We also propose an initial-guess-free algorithm to avoid its initialization dependency. Moreover, we present two global optimization techniques. The experimental results exhibit significant clustering performances from our proposed methods, including 100% clustering accuracy in the COIL-20 dataset. Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 2 |
| 2018 | Fast and Robust Estimation for Unit-Norm Constrained Linear Fitting ProblemsabstractM-estimator using iteratively reweighted least squares (IRLS) is one of the best-known methods for robust estimation. However, IRLS is ineffective for robust unit-norm constrained linear fitting (UCLF) problems, such as fundamental matrix estimation because of a poor initial solution. We overcome this problem by developing a novel objective function and its optimization, named iteratively reweighted eigenvalues minimization (IREM). IREM is guaranteed to decrease the objective function and achieves fast convergence and high robustness. In robust fundamental matrix estimation, IREM performs approximately 5-500 times faster than random sampling consensus (RANSAC) while preserving comparable or superior robustness. Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 2 |
| 2018 | Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain AdaptationabstractCan we detect common objects in a variety of image domains without instance-level annotations? In this paper, we present a framework for a novel task, cross-domain weakly supervised object detection, which addresses this question. For this paper, we have access to images with instance-level annotations in a source domain (e.g., natural image) and images with image-level annotations in a target domain (e.g., watercolor). In addition, the classes to be detected in the target domain are all or a subset of those in the source domain. Starting from a fully supervised object detector, which is pre-trained on the source domain, we propose a two-step progressive domain adaptation technique by fine-tuning the detector on two types of artificially and automatically generated samples. We test our methods on our newly collected datasets containing three image domains, and achieve an improvement of approximately 5 to 20 percentage points in terms of mean average precision (mAP) compared to the best-performing baselines. Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 3 |
| 2018 | Joint Optimization Framework for Learning With Noisy LabelsabstractDeep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can correct labels during training by alternating update of network parameters and labels. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods. Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 3 |
| 2018 | Efficient and Interactive Spatial-Semantic Image Retrieval
Ryosuke Furuta, Naoto Inoue, Toshihiko Yamasaki |
MMM (1) | 3 |
| 2018 | Efficiency-enhanced cost-volume filtering featuring coarse-to-fine strategyabstractCost-volume filtering (CVF) is one of the most widely used techniques for solving general multi-labeling problems based on a Markov random field (MRF). However it is inefficient when the label space size (i.e., the number of labels) is large. This paper presents a coarse-to-fine strategy for cost-volume filtering that efficiently and accurately addresses multi-labeling problems with a large label space size. Based on the observation that true labels at the same coordinates in images of different scales are highly correlated, we truncate unimportant labels for cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original CVF method while maintaining a comparable level of accuracy. Although we performed experiments that deal with only stereo matching and optical flow estimation, the proposed method can be employed in many other applications because of the applicability of CVF to general discrete pixel-labeling problems based on an MRF. Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 3 |
| 2018 | Photo aesthetic quality estimation using visual complexity features
Litian Sun, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 2 |
| 2018 | Fast Volume Seam Carving With Multipass Dynamic ProgrammingabstractIn volume seam carving, i.e., seam carving for 3D cost volume, an optimal seam surface can be derived by graph cuts, resulting from sophisticated graph construction. To date, the graph-cut algorithm is the only solution for volume seam carving. However, it is not suitable for practical use because it incurs a heavy computational load. We propose a multipass dynamic programming (DP)-based approach for volume seam carving, which reduces computation time and memory consumption while maintaining a similar image quality as that of graph cuts. Our multipass DP scheme is achieved by conducting DP in two directions to accumulate the cost in a 3D volume and then tracing back to find the best seam. In our multipass DP, a suboptimal seam surface is created instead of a global optimal one, and it has been experimentally confirmed by more than 198 crowdsourced workers that such suboptimal seams are good enough for image processing. The proposed scheme offers two options: a continuous method that ensures the connectivity of seam surfaces and a discontinuous method that ensures the connectivity in only one direction. We applied the proposed volume seam carving method based on multipass DP to conventional video retargeting and tone mapping. These two applications are completely different; however, the volume seam carving method can be applied similarly by changing the axes of the cost volume. Even though the results obtained using our methods were similar to those obtained by graph cuts, our computation time was approximately 90 times faster that of graph cuts and the memory usage was eight times smaller than that of graph cuts. We also extend the idea of tone mapping to the contrast enhancement method based on volume seam carving. Ryosuke Furuta, Ikuko Tsubaki, Toshihiko Yamasaki |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | PQTable: Nonexhaustive Fast Search for Product-Quantized Codes Using Hash TablesabstractIn this paper, we propose a product quantization table (PQTable)-a fast search method for product-quantized codes via hash tables. An identifier of each database vector is associated with the slot of a hash table by using its PQ-code as a key. For querying, an input vector is PQ-encoded and hashed, and the items associated with that code are then retrieved. The proposed PQTable produces the same results as a linear PQ scan, and is 102-105times faster. Although the state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods require manually designed parameter setting and significant training; our PQTable is free of these limitations, and therefore offers a practical and effective solution for real-world problems. Specifically, when the vectors are highly compressed, our PQTable achieves one of the fastest search performances on a single CPU to date with significantly efficient memory usage (0.059-ms per query over 109data points with just 5.5-GB memory consumption). Finally, we show that our proposed PQTable can naturally handle the codes of an optimized product quantization (OPQTable). Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Multim. | 2 |
| 2017 | Residual Expansion Algorithm: Fast and Effective Optimization for Nonconvex Least Squares ProblemsabstractWe propose the residual expansion (RE) algorithm: a global (or near-global) optimization method for nonconvex least squares problems. Unlike most existing nonconvex optimization techniques, the RE algorithm is not based on either stochastic or multi-point searches, therefore, it can achieve fast global optimization. Moreover, the RE algorithm is easy to implement and successful in high-dimensional optimization. The RE algorithm exhibits excellent empirical performance in terms of k-means clustering, point-set registration, optimized product quantization, and blind image deblurring. Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 2 |
| 2017 | Infrasonic scene fingerprinting for authenticating speaker locationabstractAmbient infrasound with frequency ranges well below 20 Hz is known to carry robust navigation cues that can be exploited to authenticate the location of a speaker. Unfortunately, many of the mobile devices like smartphones have been optimized to work in the human auditory range, thereby suppressing information in the infrasonic region. In this paper, we show that these ultra-low frequency cues can still be extracted from a standard smartphone recording by using acceleration-based cepstral features. To validate our claim, we have collected smartphone recordings from more than 30 different scenes and used the cues for scene fingerprinting. We report scene recognition rates in excess of 90% and a feature set analysis reveals the importance of the infrasonic signatures towards achieving the state-of-the-art recognition performance. Kenji Aono, Shantanu Chakrabartty, Toshihiko Yamasaki |
ICASSP | 3 |
| 2017 | Object detection refinement using Markov random field based pruning and learning based rescoringabstractContextual information such as the co-occurrence of objects and the location of objects has played an important role in object detection. We present candidate pruning and object rescoring methods that leverage contextual information and that can improve the state-of-the-art CNN-based object detection methods such as Fast R-CNN and Faster R-CNN. In our pruning method, we formulate candidate reduction as a Markov random field optimization problem. In our rescoring method, we employ a machine learning technique to reconsider the detection scores of candidate windows. We experimentally demonstrate improvements in R-CNN-based object detection methods using two datasets. Moreover, we apply our model to the structured retrieval task to show the potential applications of our model. Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICASSP | 3 |
| 2017 | Hyperlapse generation of omnidirectional videos by adaptive sampling based on 3D camera positionsabstractCapturing a city with an omnidirectional camera can produce a continuous motion street view and provides the view of the moving person. However, watching such a video is not necessarily pleasant because the videos are often excessively lengthy. In this paper, we propose a method for shortening an omnidirectional video. Subsampling a video captured by a hand-held camera displays a significant amount of destabilization resulting from shaking. This prompted us to propose an adaptive subsampling scheme that selects optimal frames by minimizing our cost function based on 3D camera positions. This optimal selection suppresses only the translational camera instabilities. The rotational instabilities are initially ignored and later compensated for. This approach allowed us to successfully generate a subsampled and stable omnidirectional video. In addition, we propose an alternative measure of translational instability to evaluate our frame selection. Masanori Ogawa, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2017 | How competitive are you: Analysis of people's attractiveness in an online dating systemabstractAn increasing number of people are using dating websites to search for their life partners. This leads to the curiosity of how attractive a specific person is to the opposite gender on an average level. We propose a novel algorithm to evaluate people's objective attractiveness based on their interactions with other users on the dating websites and implement machine learning algorithms to predict their objective attractiveness ratings from their profiles. We validate our method on a large dataset gained from a Japanese dating website and yield convincing results. Our prediction based on users' profiles, which includes image and text contents, is over 80% correlated with the real values of the calculated objective attractiveness for the female and over 50% correlated with the real values of the calculated objective attractiveness for the male. Xiaoxue Zang, Toshihiko Yamasaki, Kiyoharu Aizawa, Tetsuhiro Nakamoto, Eitaro Kuwabara, Shinichi Egami, Yusuke Fuchida |
ICME | 2 |
| 2017 | FolkPopularityRank: Tag Recommendation for Enhancing Social Popularity using Text Tags in Content Sharing ServicesabstractIn this study, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present the FolkPopularityRank algorithm, which can score text tags based on their ability to influence the popularity-related numbers. The FolkPopularityRank algorithm is inspired by the PageRank and FolkRank algorithms but the scores of the tags are calculated not only by the co-occurrence of the tags but also by considering the popularity-related numbers of the content. To the best of our knowledge, this is the first attempt to recommending tags that can enhance popularity attributes of social media. We conducted extensive experiments with about 1,000 images. We uploaded the photos with the recommended tags along with the original tags to Flickr as a real test, and obtained very promising results. Toshihiko Yamasaki, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa |
IJCAI | 1 |
| 2017 | Become Popular in SNS: Tag Recommendation using FolkPopularityRank to Enhance Social PopularityabstractIn this demo, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present a demo using our FolkPopularityRank (FP-Rank) algorithm, which can score and recommend text tags based on their ability to influence the popularity-related numbers. Our experiments using 1,000 photos showed that we can achieve 1.6 times more views than the original tag sets in Flickr just by adding tags recommended by FP-Rank. Toshihiko Yamasaki, Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa |
IJCAI | 1 |
| 2017 | PQk-means: Billion-scale Clustering for Product-quantized CodesabstractData clustering is a fundamental operation in data analysis. For handling large-scale data, the standard k-means clustering method is not only slow, but also memory-inefficient. We propose an efficient clustering method for billion-scale feature vectors, called PQk-means. By first compressing input vectors into short product-quantized (PQ) codes, PQk-means achieves fast and memory-efficient clustering, even for high-dimensional vectors. Similar to k-means, PQk-means repeats the assignment and update steps, both of which can be performed in the PQ-code domain. Experimental results show that even short-length (32 bit) PQ-codes can produce competitive results compared with k-means. This result is of practical importance for clustering in memory-restricted environments. Using the proposed PQk-means scheme, the clustering of one billion 128D SIFT features with K = 105 is achieved within 14 hours, using just 32 GB of memory consumption on a single computer. Yusuke Matsui 0001, Keisuke Ogaki, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 3 |
| 2017 | A Tag Recommendation System for Popularity BoostingabstractIn order to support users in the tagging process and recommendation, we had proposed two tag ranking algorithms, Document Frequency-Weights from regression and Folk Popularity Rank, which can extract tags greatly influencing popularity. We have developed a tag recommendation system using the algorithm we proposed. The recommended tags are not only for appropriate annotations but also for popularity boosting. Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 4 |
| 2017 | Sketch-based manga retrieval using manga109 datasetabstractManga (Japanese comics) are popular worldwide. However, current e-manga archives offer very limited search support, i.e., keyword-based search by title or author. To make the manga search experience more intuitive, efficient, and enjoyable, we propose a manga-specific image retrieval system. The proposed system consists of efficient margin labeling, edge orientation histogram feature description with screen tone removal, and approximate nearest-neighbor search using product quantization. For querying, the system provides a sketch-based interface. Based on the interface, two interactive reranking schemes are presented: relevance feedback and query retouch. For evaluation, we built a novel dataset of manga images, Manga109, which consists of 109 comic books of 21,142 pages drawn by professional manga artists. To the best of our knowledge, Manga109 is currently the biggest dataset of manga images available for research. Experimental results showed that the proposed framework is efficient and scalable (70 ms from 21,142 pages using a single computer with 204 MB RAM). Yusuke Matsui 0001, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 6 |
| 2016 | Uncalibrated Photometric Stereo by Stepwise Optimization Using Principal Components of Isotropic BRDFsabstractThe uncalibrated photometric stereo problem for non-Lambertian surfaces is challenging because of the large number of unknowns and its ill-posed nature stemming from unknown reflectance functions. We propose a model that represents various isotropic reflectance functions by using the principal components of items in a dataset, and formulate the uncalibrated photometric stereo as a regression problem. We then solve it by stepwise optimization utilizing principal components in order of their eigenvalues. We have also developed two techniques that lead to convergence and highly accurate reconstruction, namely (1) a coarse-to-fine approach with normal grouping, and (2) a randomized multipoint search. Our experimental results with synthetic data showed that our method significantly outperformed previous methods. We also evaluated the algorithm in terms of real image data, where it gave good reconstruction results. Keisuke Midorikawa, Toshihiko Yamasaki, Kiyoharu Aizawa |
CVPR | 2 |
| 2016 | Text detection in manga by combining connected-component-based and region-based classificationsabstractAs manga (Japanese comics) have become common content in many countries, it is necessary to search manga by text query or translate them automatically. For these applications, we must first extract texts from manga. In this paper, we develop a method to detect text regions in manga. Taking motivation from methods used in scene text detection, we propose an approach using classifiers for both connected components and regions. We have also developed a text region dataset of manga, which enables learning and detailed evaluations of methods used to detect text regions. Experiments using the dataset showed that our text detection method performs more effectively than existing methods. Yuji Aramaki, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 3 |
| 2016 | Fast volume seam carving with multi-pass dynamic programmingabstractIn volume seam carving, seam carving for three-dimensional (3D) cost volume, an optimal seam surface can be derived by graph cuts, resulting from sophisticated graph construction. However, the graph cuts algorithm is not suitable for practical use because it incurs a heavy computational load. We propose a multi-pass dynamic programming (DP) based approach for volume seam carving that reduces computation time to 60 times faster and memory consumption to 10 times smaller than those of graph cuts, while maintaining a similar image quality as that of graph cuts. In our multi-pass DP, a suboptimal seam surface is created instead of a globally optimal one, but it has been experimentally confirmed by more than 198 crowd workers that such suboptimal seams are good enough for image processing. Ryosuke Furuta, Ikuko Tsubaki, Toshihiko Yamasaki |
ICIP | 3 |
| 2016 | Interactive region segmentation for mangaabstractManga (Japanese comics) are popular all over the world, and are created digitally. In this paper, we propose an interactive segmentation method tailored for manga. The proposed method enables annotators to select areas in manga efficiently. Our experimental results showed that the proposed framework works better than Adobe Photoshop CC, which is the most widely used commercial image editing software. Kota Ito, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICPR | 3 |
| 2016 | Sketch simplification by classifying strokesabstractIn this paper, we propose a novel approach to creating clean line drawing from a scribbled sketch automatically. The main problem is determining which strokes of a scribbled sketch should be merged. We use a machine learning approach to solve this problem. Our method can automatically generate training data by comparing scribbled sketches with manually drawn line drawings without using annotations. In order to verify the generated training data, we merged strokes and created clean line drawings in accordance with the generated training data. In addition, we trained a support vector machine to estimate the pairs of strokes to be merged. Further, we verified that our method can create line drawings using this estimator. Toru Ogawa, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICPR | 3 |
| 2015 | PQTable: Fast Exact Asymmetric Distance Neighbor Search for Product Quantization Using Hash TablesabstractWe propose the product quantization table (PQTable), a product quantization-based hash table that is fast and requires neither parameter tuning nor training steps. The PQTable produces exactly the same results as a linear PQ search, and is 102 to 105 times faster when tested on the SIFT1B data. In addition, although state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods do require manually designed parameter setting and much training, whereas our method is free from them. Therefore, PQTable offers a practical and useful solution for real-world problems. Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICCV | 2 |
| 2015 | Searching for nearest neighbors with a dense space partitioningabstractProduct quantization based approximate nearest neighbor search with the use of inverted index structures have recently received increasing attention. In this paper, we propose a new inverted index structure for searching nearest neighbors in very large datasets of high dimensional data. For data indexing, our proposed method creates a dense space partitioning using multiple centroids based assigning, which generates shorter candidate lists and improves the search speed. Our experiments with a dataset of one billion SIFT features show that while achieving higher accuracy, our method demonstrates better performances on search speed compared to IV-FADC, the conventional product quantization based inverted index structure. Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 3 |
| 2015 | Fast Face Model Reconstruction and Synthesis Using an RGB-D Camera and Its Subjective EvaluationabstractIt is difficult to show a frontal face in video chatting because there is a gap between a display and a camera. We propose a method for real-time face reorientation by creating a 2.5-D face model from a single RGB-D camera and synthesizing the rotated face model with the original face image. Our method uses two kinds face models complementarily: a point cloud based model and a generic face model fitted to the user. We conducted subjective evaluation and confirmed the validity of our proposed system. Toshihiko Yamasaki, Ibuki Nakamura, Kiyoharu Aizawa |
ISM | 1 |
| 2015 | Selective K-means Tree SearchabstractIn object recognition and image retrieval, an inverted indexing method is used to solve the approximate nearest neighbor search problem. In these tasks, inverted indexing provides a nonexhaustive solution to large-scale search. However, a problem of previous inverted indexing methods is that a large-scale inverted index is required to achieve a high search recall rate. In this study, we address the problem of reducing the time required to build an inverted index without degrading the search accuracy and speed. Thus, we propose a selective k-means tree search method that combines the power of both hierarchical k-means tree and selective nonexhaustive search. Experiments based on approximate nearest neighbor search using a large dataset comprising one billion SIFT features showed that the hierarchical inverted file based on the selective k-means tree method could be built six times faster, while obtaining almost the same recall and search speed as the state-of-the-art inverted indexing methods. Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 3 |
| 2014 | Coarse-to-fine strategy for efficient cost-volume filteringabstractCost-volume filtering is one of the most widely known techniques to solve general multi-label problems, however it is problematically inefficient when the label space size is extremely large. This paper presents a coarse-to-fine strategy of the cost-volume filtering that handles efficiently and accurately multi-label problems with a large label space size. Based upon the observation that true labels at the same image coordinate of different scales are highly correlated, we truncate unimportant labels for the cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original cost-volume filtering while enjoying the comparable accuracy to it. Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 3 |
| 2014 | Multi-stage object classification featuring confidence analysis of classifier and inclined local Naive Bayes nearest neighborabstractWe propose a two-stage classification framework for image recognition which conjunctively uses parametric and non-parametric approaches. In the first stage, input images are classified using a bag-of-features (BoF) based method with a multi-class classifier. The results are categorized into two groups by our confidence analysis: highly confident and less confident. The images with less confidence are re-classified in the second stage using our inclined local naive bayes nearest neighbor (IL-NBNN). In the original local NBNN, the similarity between the input image and its k-NN classes are calculated aiming at higher discriminability and computational efficiency. Our IL-NBNN virtually calculates the similarity between all the classes efficiently by incorporating the confidence order obtained in the first stage. As a result, efficient and accurate image classification has been achieved with very small extra cost. The experiments using the three image datasets show the validity of our proposed algorithm. Takaki Maeda, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2014 | Degree of loop assessment in microvideoabstractThis paper presents a degree-of-loop assessment method for microvideo clips. Loop video is one of the popular features in microvideo, but there are so many non-loop video tagged with “loop” on microvideo services. This is because upload-ers or spammers also know that loop video is popular and they want to draw attention from viewers. In this paper, we statistically analyze the scene dynamics of the video by using color, optical flow, saliency maps, and evaluate the degree-of-loop. We have collected more than 1,000 video clips from Vine and subjectively evaluated their degree-of-loop. Experimental results show that our proposed algorithm can classify loop/non-loop video with 85.7% accuracy and categorize them into five degree-of-loop categories with 61.5% accuracy. Shumpei Sano, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2014 | Emerging Topics on Personalized and Localized Multimedia Information SystemsabstractWe are experiencing an era with a rapid increase of data relevant to different aspects of users' daily life. On the one hand, such data contains personal information of each individual user. On the other hand, it also reflects user behaviors related to the society as data of more users is aggregated. These data could not only be very beneficial for studying various lifestyle patterns, but also be used to generate more descriptive and explanatory analysis across the landscape of diverse multimedia data. Using personal mobile devices and web services to systematically explore interesting aspects of people world has attracted much attention recently. This is a full-day tutorial that addresses emerging topics on personalized and localized multimedia technologies and applications and emphasizes knowledge sensing and discovery in multimedia landscape. This tutorial aims to deliver anoverall introduction to multimedia landscapes with multimedia processing, contextual data acquisition, people activity logs, data analytics, geographic-aware multimedia sharing and delivery, and serves as an important lecture on fundamental and advanced research areas of personalized and localized multimedia information systems. Yi Yu 0001, Kiyoharu Aizawa, Toshihiko Yamasaki, Roger Zimmermann |
ACM Multimedia | 3 |
| 2014 | SVM is not always confident: Telling whether the output from multiclass SVM is true or false by analysing its confidence valuesabstractThis paper presents an algorithm to distinguish whether the output label that is yielded from multiclass support vector machine (SVM) is true or false without knowing the answer. Such judgment is done only by the confidence analysis based on the pre-training/testing using the training data. Such true/false judgment is useful for refining the output labels. We experimentally demonstrate that the decision value difference between the top candidate and the second candidate is a good measure. In addition, a proper threshold can be determined by the pre-training/testing using only the training data. Experimental results using three standard image datasets demonstrate that our proposed algorithm can improve Matthews correlation coefficient (MCC) much better than simply thresholding the decision value for the top candidate. Toshihiko Yamasaki, Takaki Maeda, Kiyoharu Aizawa |
MMSP | 1 |
| 2013 | Cooperative estimation of human motion and surfaces using multiview videos
Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa |
Comput. Vis. Image Underst. | 2 |
| 2012 | Relative-distance-based soft voting for feature representation and its application to human attribute analysisabstractThis paper proposes a soft voting based bag-of-features (BoF) model considering relative distance of the feature vectors to the nearest-neighbor codeword. Whereas state-of-the-art kernel distance based soft voting methods require brute force parameter optimization, which is time consuming, the proposed method does not require any optimization. The proposed algorithm was applied to human attribute analysis using top-view images. The experimental results have demonstrated 100% of accuracy for both gender classification and baggage possession classification. It has also been demonstrated that discriminative ability is comparable to that of the fine-tuned codeword uncertainty (UNC) model. Toshihiko Yamasaki |
ICASSP | 1 |
| 2012 | Face Recognition Challenge: Object Recognition Approaches for Human/Avatar ClassificationabstractRecently, a novel "completely automated public Turing test to tell computers and humans apart (CAPTCHA)'' system has been proposed, in which users are asked to separate natural faces of humans and artificial faces of virtual world avatars. The system is based on the assumption that computers cannot separate them while it is an easy task for humans. Conventional digital forensics approaches to distinguish natural images from computer graphics images are mostly based on statistical analysis of the images such as noise in CMOS image sensors or Bayer matrix estimation. On the other hand, this paper uses face recognition and object classification based approaches. The experiments show that our approaches work surprisingly well and yields more than 99% accuracy. Our object classification based approach can also tell us how likely the input images are regarded as human/avatar faces. Toshihiko Yamasaki, Tsuhan Chen |
ICMLA (2) | 1 |
| 2012 | Confidence-assisted classification result refinement for object recognition featuring TopN-Exemplar-SVM
Toshihiko Yamasaki, Tsuhan Chen |
ICPR | 1 |
| 2012 | Pedestrian Attribute Analysis Using a Top-View Camera in a Public Space
Toshihiko Yamasaki, Tomoaki Matsunami |
MMM | 1 |
| 2012 | Real-time tracking of humans and visualization of their future footsteps in public indoor environments - An intelligent interactive system for public entertainmentabstractIn this work, an interactive entertainment system which employs multiple-human tracking from a single camera is presented. The proposed system robustly tracks people in an indoor environment and displays their predicted future footsteps in front of them in real-time. The system is composed of a video camera, a computer and a projector. There are three main modules: tracking, analysis and visualization. The tracking module extracts people as moving blobs by using an adaptive background subtraction algorithm. Then, the location and orientation of their next footsteps are predicted. The future footsteps are visualized by a high-paced continuous display of foot images in the predicted location to simulate the natural stepping of a person. To evaluate the performance, the proposed system was exhibited during a public art exhibition in an airport. People showed surprise, excitement, curiosity. They tried to control the display of the footsteps by making various movements. Ovgu Ozturk, Tomoaki Matsunami, Yasuhiro Suzuki, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 4 |
| 2012 | Determination of emotional content of video clips by low-level audiovisual features - A dimensional and categorial experimental approachabstractAffective analysis of video content has greatly increased the possibilities of the way we perceive and deal with media. Different kinds of strategies have been tried, but results are still opened to improvements. Most of the problems come from the lack of standardized test set and real affective models. In order to cope with these issues, in this paper we describe the results of our work on the determination of affective models for evaluation of video clips using audiovisual low-level features. The affective models were developed following two classes of psychological theories of affect: categorial and dimensional. The affective models were created from real data, acquired through a series of user experiments. They reflect the affective state of a viewer after watching a certain scene from a movie. We evaluate the detection of Pleasure, Arousal and Dominance coefficients as well as the detection rate of six affective categories. For this end, two Bayesian network topologies are used, a Hidden Markov Model and an Autoregressive Hidden Markov Model. The measurements were done using audio-only models, video-only models and fused models. Fusion is done using two different methods, a Decision Level Fusion and Feature Level Fusion. All tests were conducted using localized affective models, both categorial and dimensional. Results are presented in terms of detection rate and accuracy for affective families, affective dimensions and probabilistic networks. Arousal was the best detected dimension, followed by dominance and pleasure. René Marcelino Abritta Teixeira, Toshihiko Yamasaki, Kiyoharu Aizawa |
Multim. Tools Appl. | 2 |
| 2011 | Robust watermark extraction using SVD-based dynamic stochastic resonanceabstractIn this paper, a novel dynamic stochastic resonance (DSR)-based non-blind watermark extraction technique has been proposed for robust extraction of a grayscale watermark. The watermark embedding has been carried out using singular value decomposition (SVD). Dynamic stochastic resonance has been strategically used to improve the robustness of the extraction algorithm by utilizing the noise added during attacks itself. Resilience of this technique to attacks has been tested in the presence of various noise, geometrical, enhancement, compression and filtering attacks. Using the DSR-based proposed extraction algorithm, a very robust extraction of watermark can be done without trading-off with visual quality of the watermarked image. Performance of the proposed technique has also been compared with the plain SVD-based and hybrid DCT-SVD based technique and is found to give better performance. Rajlaxmi Chouhan, Rajib Kumar Jha, Apoorv Chaturvedi, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 4 |
| 2011 | Marker-less human pose estimation and surface reconstruction using a segmented modelabstractWe propose a human motion tracking method for fast motion clips using synchronized multiple cameras. Our method is capable of extracting 3D articulated postures with 42 degrees of freedom through a sequence of visual hulls. We seek for the globally optimal solutions of the likelihood with the lo cal memorization about the "fitness" of each body segment. Our method avoids the local minimum problem efficiently by mean combination and articulated combination of parti cles selected based on the weights of the different body seg ments. We deform the template surface model using the mo tion tracking data by linear blend skinning (LBS). Then we recover the details of the surface by fitting the deformed sur face to 2D silhouettes. Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2010 | Automatic preview video generation for mesh sequencesabstractWe present a novel method that automatically generates a preview video of a mesh sequence. To make the preview appealing to users, the important features of the mesh model should be captured in the preview video, while preserving the constraint that the transitions of the camera are as smooth as possible. Our approach models the important features by defining a surface saliency and by measuring the appearance of the mesh sequence. The task of generating the preview video is then formulated as a shortest-path problem and we find an optimal camera path by using Dijkstra's algorithm. Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2010 | Patch-based compression for Time-Varying MeshesabstractThis paper proposes intra-frame and inter-frame coding algorithms for 3D mesh sequences generated by multiple cameras, which we call Time-Varying Meshes (TVMs). For this purpose, mesh segmentation into patches with patch alignment using principal component analysis (PCA) is proposed. The patches are used as minimum units to eliminate spatial and temporal correlation of the TVMs. For intraframe coding of geometry data, spectral compression using the Kirchhoff matrix was employed and that of color texture was conducted using vector quantization (VQ). The inter-frame coding of geometry data and color texture was achieved by the combination of patch-based matching and simple scalar quantization. The intra- and inter-frame coding are compared to the previous works and demonstrated promising results. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 1 |
| 2010 | Image processing based approach to food balance analysis for personal food loggingabstractFood images have been receiving increased attention in recent dietary control methods. We present the current status of our web-based system that can be used as a dietary management support system by ordinary Internet users. The system analyzes image archives of the user to identify images of meals. Further image analysis determines the nutritional composition of these meals and stores the data to form a Foodlog. The user can view the data in different formats, and also edit the data to correct any mistakes that occurred during image analysis. This paper presents detailed analysis of the performance of the current system and proposes an improvement of analysis by pre-classification and personalization. As a result, the accuracy of food balance estimation is significantly improved. Keigo Kitamura, Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 3 |
| 2010 | Detecting Dominant Motion Flows in Unstructured/Structured Crowd ScenesabstractDetecting dominant motion flows in crowd scenes is one of the major problems in video surveillance. This is particularly difficult in unstructured crowd scenes, where the participants move randomly in various directions. This paper presents a novel method which utilizes SIFT features' flow vectors to calculate the dominant motion flows in both unstructured and structured crowd scenes. SIFT features can represent the characteristic parts of objects, allowing robust tracking under non-rigid motion. First, flow vectors of SIFT features are calculated at certain intervals to form a motion flow map of the video. Next, this map is divided into equally sized square regions and in each region dominant motion flows are estimated by clustering the flow vectors. Then, local dominant motion flows are combined to obtain the global dominant motion flows. Experimental results demonstrate the successful application of the proposed method to challenging real-world scenes. Ovgu Ozturk, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICPR | 2 |
| 2010 | Automatic trailer generationabstractThis paper presents a content-based movie trailer generation method, named Vid2Trailer (V2T). Since trailers are intended to advertise movies, they must show specific symbols such as the title logo and the main theme music. Moreover, it is expected to attract viewers by its visual and audio content. V2T satisfies these two requirements when creating a trailer from the original movie content. First, the title logo and the main theme music are extracted. Second, impressive speech and video segments are extracted by using an affective content analysis technique. Third, all of the extracted components are concatenated into the form of a trailer; to realize this, we propose a method that estimates the affective impact of shot sequences, and introduce an algorithm that arranges a set of shots so as to maximize the affective impact of the sequence. Experiments show that our V2T is more appropriate to trailer generation than conventional techniques. Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 4 |
| 2010 | 3D pose estimation in high dimensional search spaces with local memorizationabstractIn this paper, a stochastic approach for extracting the articulated 3D human postures by synchronized multiple cameras in the high-dimensional configuration spaces is presented. Annealed Particle Filtering (APF) [1] seeks for the globally optimal solution of the likelihood. We improve and extend the APF with local memorization to estimate the suited kinematic postures for a volume sequence directly instead of projecting a rough simplified body model to 2D images. Our method guides the particles to the global optimization on the basis of local constraints. A segmentation algorithm is performed on the volumetric models and the process is repeated. We assign the articulated models 42 degrees of freedom. The matching error is about 6% on average while tracking the posture between two neighboring frames. Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa |
PCS | 2 |
| 2010 | Bit allocation of vertices and colors for patch-based coding in time-varying meshesabstractThis paper discusses bit-rate assignments for vertices, color, reference frames, and target frames in the patch-based compression method for time-varying meshes (TVMs). TVMs are nonisomorphic 3D mesh sequences of the real-world objects generated from multiview images. Experimental results demonstrate that the bit rate for vertices greatly affects the visual quality of the rendered 3D model, whereas the bit rate for color does not contribute to quality improvement. Therefore, as many bits as possible should be assigned to vertices, with 8–10 bits per vertex (bpv) per frame being sufficient for color. For interframe coding, the visual quality is improved in proportion to the bit rate of both vertices and color. However, it is demonstrated that the use of fewer bits (5∼6 bpv) is sufficient to achieve a visual quality that matches the intraframe visual quality. Toshihiko Yamasaki, Kiyoharu Aizawa |
PCS | 1 |
| 2010 | Approaches to 3D video compressionabstractThree-dimensional (3-D) video provides an immersing experience for users. In recent years, many attempts have been made to capture the complex surface shape and highly detailed texture of real-world moving objects, which results in a huge amount of data. In this paper, we discuss compression issues for 3-D video. We introduce 3-D video, which is classified into two categories. Then we survey compression methods that have been investigated for each category. We present our compression methods for temporally varying mesh sequences. In addition, we show comparison results for our algorithm with respect to previous work. Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa |
VCIP | 2 |
| 2010 | Affective Audio-Visual Words and Latent Topic Driving Model for Realizing Movie Affective Scene ClassificationabstractThis paper presents a novel method for movie affective scene classification that outputs the emotion (in the form of labels) that the scene is likely to arouse in viewers. Since the affective preferences of users play an important role in movie selection, affective scene classification has the potential to develop more attractive user-centric movie search and browsing applications. Two main issues in designing movie affective scene classification are considered. One is “how to extract features that are strongly related to the viewer's emotions”, and the other is “how to map the extracted features to the emotion categories”. For the former, we propose a method to extract emotion-category-specific audio-visual features named affective audio-visual words (AAVWs). For the latter issue, we propose a classification model named latent topic driving model (LTDM). Assuming that viewers' emotions are dynamically changed by the movie scene sequences, LTDM models emotions as Markovian dynamic systems driven by the sequential stimuli of the movie content. Experiments on 206 movie scenes extracted from 24 movie titles and the corresponding labels of eight emotion categories given by 16 subjects show that our method outperforms conventional approaches in terms of the subject agreement rate. Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Multim. | 4 |
| 2009 | An object-based non-blind watermarking that is robust to non-linear geometrical distortion attacksabstractThis paper presents an object-based non-blind watermarking technique that is robust to non-linear geometrical distortion attacks. This has been one of the most challenging problems for copyright protection of digital content because it is difficult to estimate the distortion parameters for the embedded blocks. In our proposed scheme, the locations used to embed the watermark information are memorized and detected robustly by using relative coordinates with respect to the scale-invariant feature transform (SIFT) feature points. Experimental results with 64-bit watermark embedding demonstrated that the watermark detection performance was improved from 57% to 82% on average, even after nonlinear geometrical attacks such as waving, imploding, swirling, and seam carving. Toshihiko Yamasaki, Yasumasa Nakai, Kiyoharu Aizawa |
ICIP | 1 |
| 2009 | Affective video segment retrieval for consumer generated videos based on correlation between emotions and emotional audio eventsabstractA novel affective video segment retrieval method based on the correlation between emotion and emotional audio events (EAEs) is presented. The proposed method focuses on retrieving three types of affective video segments, joy, sadness and excitement, by utilizing correlations between emotions and EAEs. The correlation between these emotions and EAEs is investigated by a subjective evaluation. The proposed method detects EAEs and rates each EAE in terms of emotion levels. The EAEs are detected by using the generalized state-space model (GSSM) and low-level audio features. Experiments conducted on consumer generated videos (CGVs) show that the proposed EAE detection outperforms conventional HMM and GMM based methods in terms of accuracy, the agreement rate of the retrieved affective video segments reaches 73.3%. Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 4 |
| 2009 | A degree-of-edit ranking for consumer generated video retrievalabstractWe introduce degree-of-edit (DoE) ranking to focus on ldquohow much a CGV is editedrdquo as a ranking measure for consumer generated video (CGV) retrieval; a method to estimate DoE ranking is proposed. In the proposed method, the DoE score of a CGV is estimated by using low-level features such as the number of shot boundaries and time ratio of music. We evaluate the rank correlation between DoE ranking determined by subjects and by our method. To demonstrate its performance in a practical scenario, a user test is performed on over 22,000 CGVs in the context of CGV search. The obtained results show that our method significantly improves conventional CGV ranking results in terms of availabilities of interesting and high-quality CGVs. Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 4 |
| 2009 | Retrieval of Time-Varying Mesh and motion capture data using 2D video queries based on silhouette shape descriptorsabstractThis paper presents a retrieval system for time-varying mesh (TVM) and motion capture data using 2D video queries. Previous approaches have used other TVM and motion capture data as queries and the cost for query generation was a significant issue. Instead, the proposed system uses 2D video queries, which can be easily captured by a single camera, enabling end users to retrieve 3D motion sequences such as TVM and motion capture data easily and interactively. We introduce the P-type Fourier descriptor, which is a feature of 2D contour images. TVM and computer graphics sequences rendered from motion capture data are silhouetted rendering from multiple viewpoints. Feature vectors for TVM and motion capture data are generated by applying the P-type Fourier descriptor to these silhouetted images. Experimental results using four TVM sequences and motion capture data demonstrated an average retrieval accuracy of 88% in terms of nearest neighbors. Daisuke Kasai, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2009 | A Euclidean-geodesic shape distribution for retrieval of time-varying mesh sequencesabstractThis paper proposes a Euclidean-geodesic shape distribution for the more accurate retrieval of time-varying meshes, which are 3D mesh sequences of real-world objects generated by multiple cameras. The Euclidean-geodesic shape distribution derives from a combination of the modified shape distribution algorithm, which analyzes the global shape features of 3D models, and the geodesic shape distribution algorithm, which is used to investigate topological changes. The optimal weighting for the two algorithms is investigated experimentally. Experimental results show that the performance for similar motion retrieval is better than that of conventional algorithms, being improved by 2% on average and by 5% in the best case. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 1 |
| 2009 | Latent topic driving model for movie affective scene classificationabstractThis paper proposes a latent topic driving model (LTDM) as a novel approach to movie affective scene classification. LTDM is a discriminative model of emotions driven by movie affective contents. Unlike existing methods, our approach is based on movie topic extraction via the latent Dirichlet allocation (LDA) and emotion dynamics modeling with reference to Plutchik's emotion theory. The classification procedure starts by segmenting movie scenes into movie shots, each of which is represented by a histogram of quantized affect-related audio-visual features. LDA is applied to detect topics of each movie shot. Emotions for the current movie shot are estimated based on both the topics of the shot and emotion transition weights determined by Plutchik's emotion theory. We conduct experiments using 206 movie scenes extracted from 24 movie titles (total 6 hours 20 min. 12 sec.) and the labels of eight emotion categories given by 16 subjects are collected. The results show that LTDM outperforms conventional modeling approaches in terms of the subject agreement rate. Go Irie, Kota Hidaka, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 5 |
| 2009 | Sketch-on-Map: Spatial Queries for Retrieving Human Locomotion Patterns from Continuously Archived GPS Data
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
MMM | 2 |
| 2009 | A low-power switched-current CDMA matched filter employing MOS linear matching cell with on-chip A/D converter
Toshihiko Yamasaki, Tomoyuki Nakayama, Tadashi Shibata |
Integr. | 1 |
| 2009 | Temporal Segmentation of 3-D Video by Histogram-Based Feature VectorsabstractThree-dimensional (3-D) video, which is a sequence of time-varying mesh models generated in a multi-camera studio, is attracting increased attention, because it can record and reproduce the 3-D information of real-world objects with high accuracy. As one of the most important preprocessings for indexing, annotation, retrieval, and many other functions in management of a 3-D video database, it is necessary to temporally segment 3-D video into meaningful and manageable segments. We have developed robust and effective segmentation algorithms using histogram-based feature vector representation, striving to understand and manage 3-D video contents. We have developed two approaches to generate feature vectors by vertex positions in the mesh models: one uses the Cartesian coordinate system and the other employs the spherical coordinate system. Then, 3-D video is segmented by the motion intensity of an object, which is analyzed by the feature vectors. The segmentation algorithms we have developed are applied to three different 3-D video sequences. A statistical method is presented to evaluate the segmentation results. High recall and precision rates of 0.95 and 0.77, respectively, are achieved in the best case. Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Sketch-Based Spatial Queries for Retrieving Human Locomotion Patterns From Continuously Archived GPS DataabstractWe propose a system for retrieving human locomotion patterns from tracking data captured within a large geographical area, over a long period of time. A GPS receiver continuously captures data regarding the location of the person carrying it. A constrained agglomerative hierarchical clustering algorithm segments these data according to the person's navigational behavior. Sketches made on a map displayed on a computer screen are used for specifying queries regarding locomotion patterns. Two basic sketch primitives, selected based on a user study, are combined to form five different types of queries. We implement algorithms to analyze a sketch made by a user, identify the query, and retrieve results from the collection of data. A graphical user interface combines the user interaction strategy and algorithms, and allows hierarchical querying and visualization of intermediate results. We evaluate the system using a collection of data captured during nine months. The constrained hierarchical clustering algorithm is able to segment GPS data at an overall accuracy of 94% despite the presence of location-dependent noise. A user study was conducted to evaluate the proposed user interaction strategy and the usability of the overall system. The results of this study demonstrate that the proposed user interaction strategy facilitates fast querying, and efficient and accurate retrieval, in an intuitive manner. Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Multim. | 2 |
| 2008 | Error analysis of 3Dc-based normal map compression and its application to optimized quantizationabstractNormal mapping is one of the most essential technologies for realistic three-dimensional computer graphics. In conventional normal map compression such as 3Dc, only the x and y components are encoded and the z components are restored based on the normalizing condition. In this paper, we present an intuitively comprehensive error analysis for this approach. As a result, we reveal in what condition compression error becomes larger. We also present a non-linear quantization algorithm based on the formula for better compression performance than the conventional approaches. Experimental results using 300 normal map demonstrate that the PSNR is improved by 0.29 dB on average. Our algorithm is compatible with random access and highly-parallel processing on GPU. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICASSP | 1 |
| 2008 | Geometry compression for time-varying meshes using coarse and fine levels of quantization and run-length encodingabstractTime-varying meshes (TVM) is a new 3-D scene representation which are generated from multiple cameras. It captures highly detailed shape and texture as well as movement of real-world moving objects. Compression is a key technology for supporting TVM applications such as education, interactive broadcasting, and intangible heritage archiving. Previous works focused on compression of 3-D animation that has the same topology throughout the entire sequences. Unfortunately, the topology of TVMs change with time which makes it difficult to compress TVMs. In this paper, we propose a geometry encoder for TVMs. The encoder finds spatial and temporal redundancy by coarse and fine level quantization. Thereafter, vertex information is converted into binary sequences. And then, the binary sequences are encoded using run-length encoding (RLE). Experimental results show that vertices of TVMs which require 96 bits per vertex (bpv) are compressed to 1.9-15.4 bpv while maintaining a small geometric distortion ranging from 0.7 times 10-4to 1.3 times 10-3% of the maximum error. Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2008 | Hierarchical mesh decomposition and motion tracking for Time-Varying-MeshesabstractThis paper proposes a system for automatic segmentation and motion tracking of Time-Varying-Meshes (TVM). Our approach is based on skeleton-based hierarchical mesh decomposition by distance calculation. The properties of the human skeleton structure are used to define the decomposition of each TVM frame. The proposed framework is a recursive system that iterates between automatic hierarchical decomposition on minimum distance satisfaction and skeleton realignment. This is done to achieve a stable segmentation and a refined skeleton. By utilizing color information, ill-defined meshes can be successfully segmented. Results show an average disparity of 1.39% of total surface area across all segmented parts, which indicate stability of the system. In addition, motion tracking is successfully performed through the use of refined skeletons. Ning Sung Lee, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2008 | High level activity annotation of daily experiences by a combination of a wearable device and Wi-Fi based positioning systemabstractMany people would like to record and manipulate their experiences effectively. However, efficient summarization to show “what, when and where” we did in our daily lives is still an open issue in life log applications. In conventional approaches, many sensors were attached to a human body to solve this problem. However, this is not practical for daily use. In this paper, we propose a simple solution in which a user wears two devices: a single life-logging device SenseCam [1] hung by neck and a Wi-Fi enabled PDA. The location data and low-level activity are analyzed by a Wi-Fi based positioning system. Then, high-level activity classification is conducted using the wearable device along with a refinement process considering the time consistency. Experimental results demonstrated high recall and precision rates as much as 81% and 85% respectively, on average. Wayhit Puangpakisiri, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2008 | Food log by analyzing food imagesabstractIn this paper, a food-logging system that can distinguish food images from other images, analyze the food balance, and visualize the log is presented. The image processing is based on feature vectors consisting of color histograms, DCT coefficients, detected image patterns and so forth. Support Vector Machine (SVM) was used to detect food images and to analyze the food balance. Experimental results show that the food image extraction presents above 88% of accuracy and the food balance estimation is achieved with more than 73% of accuracy. Keigo Kitamura, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2008 | Interactive retrieval for multi-camera surveillance systems featuring spatio-temporal summarizationabstractAn interactive interface is presented for near-synchronized and distributed multi-camera surveillance systems. Human tracking using multiple cameras is conducted employing a particle filter in conjunction with linear least-square trajectory extrapolation and color histogram matching. Spatial and temporal statistical information such as histograms of the number of pedestrians on a timeline-basis, pedestrians' trajectories over a certain period, pedestrian flow, and so forth is efficiently summarized and visualized. In addition, a sketch-based retrieval interface is also developed. As a result, our system facilitates operators to easily extract and access to important scenes from a huge amount of surveillance data. The real-life experiments in the public street demonstrated the validity of our system. Toshihiko Yamasaki, Yoshifumi Nishioka, Kiyoharu Aizawa |
ACM Multimedia | 1 |
| 2008 | Audio Analysis for Multimedia Retrieval from a Ubiquitous Home
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
MMM | 2 |
| 2007 | A Sensor Network for Event Retrieval in a Home Like Ubiquitous EnvironmentabstractWe present the current status of a system based on a network of sensors for event retrieval from a home like environment. A large number of cameras, microphones and pressure based floor sensors are used for continuous data capture. The data are analyzed independently and the results recorded in a central database, where they are combined for efficient retrieval and summarization of the video and audio data. We describe the detection of basic actions, events, and faces using image analysis. The users can query the system interactively to retrieve video, audio, and key frames corresponding to events. We report results of performance evaluations of the algorithms and discuss issues related to sensing, analysis and retrieval of data. Gamhewage Chaminda de Silva, Steve Anavi, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICASSP (4) | 3 |
| 2007 | Highly Efficient VQ-Based Normal Map Compression using Quality Estimation ModelabstractNormal maps play an important role in computer 3D graphics to express pseudo roughness of the surface with a small amount of polygon data. In this paper, a highly efficient normal map compression algorithm is proposed based on an estimation model to predict the quality of the images rendered with the compressed normal maps. The optimal encoding is achieved by minimizing the predicted mean square error (MSE) employing vector quantization (VQ). In addition, encoding and decoding time is fast enough for practical usage. Experimental results demonstrate that the algorithm proposed in this paper yields better compression performance than the other algorithms in the literatures. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICASSP (1) | 1 |
| 2007 | Tracking Persons using Particle Filter Fusing Visual and Wi-Fi Localizations for Widely Distributed CameraabstractThis paper describes an object tracking scheme employs sensor fusion approach which is composed of visual information and location information estimated from Wi-Fi signals. Location information is calculated by a set of received signal strength values of beacon packets from Wi-Fi access points (APs) around the targets. Different from the conventional approaches which use another kind of sensors, our approach can cover wider areas both indoor and outdoor with lower cost because of characteristics of Wi-Fi signals. Particle filter is applied to combine these two different kinds of sensory input to track the target continuously. Wi-Fi observation model is involved in a conventional visual particle filtering scheme in order to evaluate importance weights of each particle. By using multiple modality, robust tracking performance is achieved even if reliability of one sensory input declines. In this paper, we present experimental results applied to outdoor surveillance camera environment. Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP (3) | 2 |
| 2007 | Geometrically Invariant Object-Based Watermarking using SIFT FeatureabstractIn this paper, we have developed a robust object-based watermarking algorithm using the scale-invariant feature transform (SIFT) features in conjunction with a new data embedding method based on discrete cosine transform (DCT). The message is embedded in DCT spaces of randomly generated blocks in the selected object region. To recognize the object region after being distorted, its SIFT features are registered in advance. In the detection scheme, we firstly detect the object region by using feature matching. The transformation parameters are then calculated, and the message can be detected. Experimental results demonstrated that our proposed algorithm is very robust to geometrical distortions such as JPEG compression, scaling, rotation, shearing, aspect ratio change, image filtering, and so on. Viet Quoc Pham, Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP (5) | 3 |
| 2007 | View-Based Web Page Retrieval using Interactive Sketch QueryabstractWe propose a novel view-based Web page retrieval system that enables a user to search Web pages using a visual query, namely the user's freehand sketch. We believe the proposed method will suit retrieval from a set of web pages such as a user's local browsing history. The system aims to help the user revisit a particular Web page without using query words. Using color signature features and Earth-Mover's distance, the system evaluates the similarity between web pages and the user's sketch drawn via the GUI of the system. In order to accelerate the interaction, the results of the similarity evaluations are shown immediately after the user draws each stroke of the sketch, with the results being interactively reordered. Experiments using our prototype system showed that users find their target pages after only a few strokes. Experimental results for inexperienced users showed that 71% of search tasks were completed within one minute, using the prototype system. The median time for the tasks was 40 seconds. Yasuyuki Watai, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP (6) | 2 |
| 2007 | Visual Tracking of Pedestrians Jointly using Wi-Fi Location System on Distributed Camera NetworkabstractObject tracking with multiple cameras is a fundamental problem in wide-area surveillance application, but it has difficulties to achieve accurate and stable performance because of disjoint shot areas or initial object identification problems. We propose a novel object tracking method which jointly uses estimated location information of the target derived from a set of Wi-Fi signal strength values with video images from cameras. Apart from sensor fusion techniques proposed in the past which use another kind of sensors (eg., global positioning system (GPS), pressure sensors on the floors to detect foot steps, laser-range scanners, etc), our approach can cover wide-areas both indoor and outdoor in low cost because of propagation characteristics of Wi-Fi signals. Particle filtering is applied to achieve tracking the target from video images, with this Wi-Fi location estimation. This paper describes system architectures and the experimental results of the pedestrian tracking technique with Wi-Fi location estimation. Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2007 | Fast and Robust Motion Tracking for Time-Varying Mesh Featuring Reeb-Graph-Based Skeleton Fitting and its Application to Motion RetrievalabstractIn this paper, an algorithm for motion extraction from time-varying mesh (TVM) is proposed. TVM is a sequence of 3D models made for real-world dynamic objects. In TVM, detailed information of 3D objects such as shape, color, and motion is stored. Therefore, TVM is an attractive technology for the next-generation multimedia. However, the management of TVM database for archiving and retrieving is a significant problem because it is difficult to locate and track feature points in TVM. For efficient and effective archive and retrieval systems, motion extraction and tracking is essential. In our approach, skeletons are extracted from TVM using Reeb graph, and motion tracking is achieved by fitting a reference skeleton to the others. We defined a geodesic function using principal component analysis (PCA) for fast and noiseless skeleton extraction. Robust fitting has been realized by an end node tracking strategy. In addition, we have developed an efficient motion retrieval system compatible to conventional motion capture (mocap) data using the extracted skeleton. As a result, low-cost query generation using mocap systems and cross search between TVM and mocap data have been achieved. Ryuichi Tadano, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2007 | Deformation of Time-Varying-Mesh Based on Semantic Human ModelabstractIn this paper, an approach is presented for deformation of time-varying-mesh, which is a sequence of 3D mesh models. The deformation here is a process to generate mid-frames between two frames that transform smoothly from one to the other. Motion vectors between two frames are extracted based on a semantic human model with an assumption of the articulated object with piecewise-rigid motions. For this purpose, each mesh model is transformed to a volumetric model, where a distance field is constructed. In the first frame, the user manually segments the volumetric model into the semantic human model. Then, fast motion estimation of the volumetric models is performed between the two frames with a cost function based on the distance field. Using the estimated motion vectors, a realistic deformation of a mesh model is achieved. Lastly, mid-frames are interpolated linearly. The technique can be applied in many areas such as frame rate up-conversion and motion blending. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2007 | Content-Based Cross Search for Human Motion Data using Time-Varying Mesh and Motion Capture DataabstractThis paper describes a content-based cross search scheme for two kinds of three-dimensional (3D) human motion data: time-varying mesh (TVM) and motion capture data. TVM is a sequence of 3D mesh models made for real-world 3D objects. TVM is generated using multiple-view images taken with multiple cameras. Since TVM can record shape, color, and motion of the real-world 3D objects, it has been drawing a lot of attention these days. In order to realize practical archiving systems for TVM, efficient retrieval systems are indispensable. The retrieval systems for TVM developed so far are based on query-by-example. This means additional TVM generation is required for constructing queries, which is computationally demanding and time consuming. On the other hand, motion capture systems are widely used to capture 3D human motion. However, the data structure is very different from that of TVM. Therefore, the two kinds of 3D human motion data are incompatible to each other. In this paper, we present a retrieval system that enables retrieving TVM using motion capture data as queries and vice versa using the modified shape distribution algorithm. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 1 |
| 2007 | Spatial querying for retrieval of locomotion patterns in smart environmentsabstractA system for retrieving video sequences created by tracking humans in a smart environment, by using spatial queries, is presented. Sketches made on a graphical user interface using a pointing device are used as the means of entering multiple types of queries. After preprocessing and coordinate system conversion, the sketches are analyzed to identify the type of the query. Directional search algorithms based on the minimum distance between points is applied for finding the best matches to the sketch. The results are ranked according to the similarity and presented to the user. The results of an initial evaluation are reported. The paper concludes with an outline of possible future directions. Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2007 | Motion Structure Parsing and Motion Editing in 3D Video
Toshihiko Yamasaki, Kiyoharu Aizawa |
MMM (1) | 2 |
| 2007 | Time-Varying Mesh Compression Using an Extended Block Matching AlgorithmabstractTime-varying mesh, which is attracting a lot of attention as a new multimedia representation method, is a sequence of 3-D models that are composed of vertices, edges, and some attribute components such as color. Among these components, vertices require large storage space. In conventional 2-D video compression algorithms, motion compensation (MC) using a block matching algorithm is frequently employed to reduce temporal redundancy between consecutive frames. However, there has been no such technology for 3-D time-varying mesh so far. Therefore, in this paper, we have developed an extended block matching algorithm (EBMA) to reduce the temporal redundancy of the geometry information in the time-varying mesh by extending the idea of the 2-D block matching algorithm to 3-D space. In our EBMA, a cubic block is used as a matching unit. MC in the 3-D space is achieved efficiently by matching the mean normal vectors calculated from partial surfaces in cubic blocks, which our experiments showed to be a suboptimal matching criterion. After MC, residuals are transformed by the discrete cosine transform, uniformly quantized, and then encoded. The extracted motion vectors are also entropy coded after differential pulse code modulation. As a result of our experiments, 10%-18% compression has been achieved. Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Fast and Efficient Normal MAP Compression Based on Vector QuantizationabstractNormal maps play an important role in realistic 3D image rendering to express pseudo roughness of the surface with small amount of polygon data. In this paper, a fast and efficient normal map compression algorithm is proposed based on vector quantization and entropy coding. Using the strong correlation among x, y, and z components of normal maps owing to the unity condition, compression ratio has been made much better than conventional approaches. In addition, the encoding time has been made reasonable by considering the distribution of the data and employing inner product in nearest-neighbor search instead of Euclidian distance taking advantage of the unity condition of the training data Toshihiko Yamasaki, Kiyoharu Aizawa |
ICASSP (2) | 1 |
| 2006 | 3D Video Compression Based on Extended Block Matching AlgorithmabstractThree dimensional (3D) video is attracting a lot of attention as a new multimedia representation method. 3D video is a sequence of 3D models (frames) that consist of varying vertices and connectivity. In conventional 2D video compression algorithms, motion compensation (MC) using block matching algorithm is frequently employed to reduce redundancy between consecutive frames. However, there is no such technology for 3D video so far. Therefore, in this paper, we have developed an extended block matching algorithm (EMBA) to reduce temporal redundancy of geometry information of 3D video by extending the idea of 2D block matching to 3D space. In our EBMA, a cubic block is used as a matching unit and, MC is achieved efficiently by matching the mean normal vectors of the sub-blocks, which turned out to be sub-optimal by our experiments. The residual information is further transformed by discrete cosine transform (DCT) and then encoded. The extracted motion vectors are also entropy encoded. As a result of our experiments, compression ratio ranging from 10% to 18% of the original 3D video data has been achieved. Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP | 2 |
| 2006 | Key Frame Extraction in 3D Video by Rate-Distortion Optimizationabstract3D video, which consists of a sequence of 3D mesh models, can provide detailed 3D information both in spatial and temporal domain. In this paper, a key frame extraction method has been developed to summarize 3D video by rate-distortion optimization. For this purpose, we introduce an effective feature vector extraction algorithm from 3D video. Prior to key frame extraction, shot detection is performed using the feature vectors as a pre-processing. Then, a rate-distortion (R-D) curve is generated in each shot, where the locations of key frames are optimized. Lastly, R-D trade-off can be achieved by optimizing a cost function with a Lagrange multiplier. Our experimental results show the extracted key frames are compact and faithful to original 3D video Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 2 |
| 2006 | Motion Segmentation of 3D Video using Modified Shape DistributionabstractIn this paper, temporal segmentation of 3D video based on motion analysis is presented. 3D video is a sequence of 3D models made for a real-world dynamic object. A modified shape distribution algorithm is proposed to realize stable shape feature representation. In our approach, representative points are generated by clustering vertices based on their spatial distribution instead of randomly sampling vertices as in the original shape distribution algorithm. Motion segmentation is conducted analyzing local minima in degree of motion calculated in the feature vector space. The segmentation algorithm developed in this paper does not require any predefined threshold values but rely on relative relationships among local minima and local maxima of the motion. Therefore, robust segmentation has been achieved. The experiments using 3D video of traditional dances yielded encouraging results with the precision and recall rates of 93% and 88%, respectively, on average Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 1 |
| 2005 | 3D video segmentation using point distance histogramsabstractSimilar to 2D video segmentation, 3D video segmentation is to divide the 3D video in temporal domain into a set of meaningful and manageable segments (shots) that are used as basic elements for indexing. This paper proposes a temporal segmentation method for 3D video for the first time as far as we know. Point distance histograms are used considering the tradeoff between computational cost and effectiveness. In order to reflect the real motion of 3D object, three fixed points are selected, which can avoid what we call "the same sphere problem." And simulation results on total 285 frames, which are composed of four different sequences, show our method is very effective. Toshihiko Yamasaki, Kiyoharu Aizawa |
ICIP (1) | 2 |
| 2005 | Mathematical error analysis of normal map compression based on unity conditionabstractNormal maps play an important role in realistic 3D image rendering to express pseudo roughness of the surface. In normal map compression, z components are often eliminated and calculated from x and y components based on the unity condition to achieve high compression rate. However, there is no theoretical background of the validity of eliminating z components so far. In this paper, a mathematical mean square error (MSE) model of such normal map compression based on the unity condition is proposed. In addition, the boundary condition whether to eliminate or include the z components in compression for the better efficient encoding is demonstrated based on our error analysis model. Toshihiko Yamasaki, Kazuya Hayase, Kiyoharu Aizawa |
ICIP (2) | 1 |
| 2005 | Video handover for retrieval in a ubiquitous environment using floor sensor dataabstractA system for retrieving video captured in a ubiquitous environment is presented. Data from pressure-based floor sensors are obtained as a supplementary input together with video from multiple stationary cameras. Unsupervised data mining techniques are used to reduce noise present in floor sensor data. An algorithm based on agglomerative hierarchical clustering is used to segment footpaths of individual persons. Video handover is proposed and two methods are implemented to retrieve video and key frame sequences showing a person moving in the house. Users can query the system based on time and retrieve video or key frames using either of the handover techniques. We compare the results of retrieval using different techniques subjectively. We conclude with suggestions for improvements, and future directions. Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Takayuki Ishikawa, Kiyoharu Aizawa |
ICME | 2 |
| 2005 | Evaluation of video summarization for a large number of cameras in ubiquitous homeabstractA system for video summarization in a ubiquitous environment is presented. Data from pressure-based floor sensors are clustered to segment footsteps of different persons. Video handover has been implemented to retrieve a continuous video showing a person moving in the environment. Several methods for extracting key frames from the resulting video sequences have been implemented, and evaluated by experiments. It was found that most of the key frames the human subjects desire to see could be retrieved using an adaptive algorithm based on camera changes and the number of footsteps within the view of the same camera. The system consists of a graphical user interface that can be used to retrieve video summaries interactively using simple queries. Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2003 | Analog soft-pattern-matching classifier using floating-gate MOS technologyabstractA flexible analog pattern-matching classifier has been developed and its performance is demonstrated in conjunction with a robust image representation algorithm called projected principal-edge distribution (PPED). In the circuit, the functional form of matching is made tunable in terms of the peak position, the peak height and the sharpness of the similarity evaluation by employing the floating-gate MOS technology. The test chip was fabricated in a 0.6-/spl mu/m complimentary metal-oxide semiconductor technology and successfully applied to the recognition of simple handwritten patterns and Arabic numerals using the PPED algorithm for robust image coding. The separation and classification of overlapping patterns have been also experimentally demonstrated. Toshihiko Yamasaki, Tadashi Shibata |
IEEE Trans. Neural Networks | 1 |
| 2001 | Analog Soft-Pattern-Matching Classifier using Floating-Gate MOS TechnologyabstractA flexible pattern-matching analog classifier is presented in con- junction with a robust image representation algorithm called Prin- cipal Axes Projection (PAP). In the circuit, the functional form of matching is configurable in terms of the peak position, the peak height and the sharpness of the similarity evaluation. The test chip was fabri- cated in a 0.6-m m CMOS technology and successfully applied to hand-written pattern recognition and medical radiograph analysis using PAP as a feature extraction pre-processing step for robust image coding. The separation and classification of overlapping patterns is also ex- perimentally demonstrated. Toshihiko Yamasaki, Tadashi Shibata |
NIPS | 1 |