Ichiro Ide

dblp:33/339 · DBLP profile ↗
← Back
95ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0003-3942-9296ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 37 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 View-aware Cross-modal Distillation for Multi-view Action Recognition
abstract
The widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
WACV5
2026 MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environments
abstract
Multi-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
Pattern Recognit.5
2026 Hierarchical Global-Local Fusion for One-stage Open-vocabulary Temporal Action Detection
abstract
Open-vocabulary Temporal Action Detection (Open-vocab TAD) extends the detection scope of Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) to unseen action classes specified by vocabularies not included in the training data, within untrimmed video. Typical Open-vocab TAD methods adopt a two-stage approach that first proposes candidate action intervals and then identifies those actions. However, errors in the first stage can affect the subsequent stage and the final detection results. Moreover, conventional methods for temporal context analyses tend to focus solely on either global or local context. Focusing solely on the global context can lead to lack of momentary detail, making it difficult to distinguish one action from another. Conversely, focusing only on the local context makes it challenging to determine the start and end timings of action intervals. To address these challenges, we introduce a one-stage approach named Hierarchical Open-vocab TAD (HOTAD), consisting of two branches: Temporal Context Analysis (TCA) and Video–Text Alignment (VTA). The former utilizes Hierarchical Encoder (HE) to fuse global and local temporal features, enabling a comprehensive capture of temporal actions, while the latter branch exploits the synergy between visual and textual modalities for precisely detecting unseen actions in the Open-vocab setting. Experiments and in-depth analysis using the widely recognized datasets THUMOS14 and ActivityNet-1.3 are performed to show the effectiveness of HOTAD. The results highlight remarkable accuracy in detecting a wide range of unseen actions. Furthermore, HOTAD significantly reduces wrong labels and localizes action instances with high precision, showcasing its robustness in complex and dynamic video settings.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
ACM Trans. Multim. Comput. Commun. Appl.4
2025 MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
abstract
Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
FG5
2025 Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
abstract
Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard “pre-training followed by fine-tuning” paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) approach offers a promising alternative, but primarily focuses on the language components of Multimodal Large Language Models (MLLMs). Despite recent progress, PEFT remains underexplored in multimodal tasks and lacks sufficient understanding of visual information during fine-tuning the model. To bridge this gap, we propose Query-Adapter (Q-Adapter), a lightweight visual adapter module designed to enhance MLLMs by enabling efficient fine-tuning for the video captioning task. Q-Adapter introduces learnable query tokens and a gating layer into Vision Encoder, enabling effective extraction of sparse, caption-relevant features without relying on external textual supervision. We evaluate Q-Adapter on two well-known video captioning datasets, MSR-VTT and MSVD, where it achieves state-of-the-art performance among the methods that take the PEFT approach across BLEU@4, METEOR, ROUGE-L, and CIDEr metrics. Q-Adapter also achieves competitive performance compared to methods that take the full fine-tuning approach while requiring only 1.4% of the parameters. We further analyze the impact of key hyperparameters and design choices on fine-tuning effectiveness, providing insights into optimization strategies for adapter-based learning. These results highlight the strong potential of Q-Adapter in balancing caption quality and parameter efficiency, demonstrating its scalability for video–language modeling.
Junan Chen 0004, Trung Thanh Nguyen 0006, Takahiro Komamizu, Ichiro Ide
MMAsia4
2025 Quantifying Image-Adjective Associations by Leveraging Large-Scale Pretrained Models
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
MMM (4)5
2025 Towards Visual Storytelling by Understanding Narrative Context Through Scene-Graphs
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (4)5
2024 R-DiP: Re-ranking Based Diffusion Pre-computation for Image Retrieval
Tatsuya Kato, Takahiro Komamizu, Ichiro Ide
DEXA (2)3
2024 One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-Scale and Action Label Features
abstract
Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing and classifying actions based on a predefined set of categories. In contrast, Open-vocab TAD goes further and is not limited to these predefined categories. This is particularly useful in real-world scenarios where the variety of actions in videos can be vast and not always predictable. The prevalent methods in Open-vocab TAD typically employ a 2-stage approach, which involves generating action proposals and then identifying those actions. However, errors made during the first stage can adversely affect the subsequent action identification accuracy. Additionally, existing studies face challenges in handling actions of different durations owing to the use of fixed temporal processing methods. Therefore, we propose a L-stage approach consisting of two primary modules: Multi-scale Video Analysis (MVA) and Video-Text Alignment (VTA). The MVA module captures actions at varying temporal resolutions, overcoming the challenge of detecting actions with diverse durations. The VTA module leverages the synergy between visual and textual modalities to precisely align video segments with corresponding action labels, a critical step for accurate action identification in Open-vocab scenarios. Evaluations on widely recognized datasets THUMOSl4 and ActivityNet-I.3, showed that the proposed method achieved superior results compared to the other methods in both Open-vocab and Closed-vocab settings. This serves as a strong demonstration of the effectiveness of the proposed method in the TAD task.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
FG4
2024 Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
abstract
Text-to-image diffusion models sometimes depict blended concepts in the generated images. One promising use case of this effect would be the nonword-to-image generation task which attempts to generate images intuitively imaginable from a non-existing word (nonword). To realize nonword-to-image generation, an existing study focused on associating nonwords with similar-sounding words. Since each nonword can have multiple similar-sounding words, generating images containing their blended concepts would increase intuitiveness, facilitating creative activities and promoting computational psycholinguistics. Nevertheless, no existing study has quantitatively evaluated this effect in either diffusion models or the nonword-to-image generation paradigm. Therefore, this paper first analyzes the conceptual blending in a pretrained diffusion model, Stable Diffusion. The analysis reveals that a high percentage of generated images depict blended concepts when inputting an embedding interpolating between the text embeddings of two text prompts referring to different concepts. Next, this paper explores the best text embedding space conversion method of an existing nonword-to-image generation framework to ensure both the occurrence of conceptual blending and image generation quality. We compare the conventional direct prediction approach with the proposed method that combines k-nearest neighbor search and linear regression. Evaluation reveals that the enhanced accuracy of the embedding space conversion by the proposed method improves the image generation quality, while the emergence of conceptual blending could be attributed mainly to the specific dimensions of the high-dimensional text embedding space.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
ACM Multimedia5
2024 Action Selection Learning for Multi-label Multi-view Action Recognition
abstract
Multi-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMAsia4
2024 Cross-modal recipe retrieval based on unified text encoder with fine-grained contrastive learning
abstract
Cross-modal recipe retrieval is vital for transforming visual food cues into actionable cooking guidance, making culinary creativity more accessible. Existing methods separately encode the recipe Title, Ingredient, and Instruction using different text encoders, then aggregate them to obtain recipe feature, and finally match it with encoded image feature in a joint embedding space. These methods perform well but require significant computational cost. In addition, they only consider matching the entire recipe and the image but ignore the fine-grained correspondence between recipe components and the image, resulting in insufficient cross-modal interaction. To this end, we propose U nified T ext E ncoder with F ine-grained C ontrastive L earning (UTE-FCL) to achieve a simple but efficient model. Specifically, in each recipe, UTE-FCL first concatenates each of the Ingredient and Instruction texts composed of multiple sentences as a single text. Then, it connects these two concatenated texts with the original single-phrase Title to obtain the concatenated recipe. Finally, it encodes these three concatenated texts and the original Title by a Transformer-based Unified Text Encoder (UTE). This proposed structure greatly reduces the memory usage and improves the feature encoding efficiency. Further, we propose fine-grained contrastive learning objectives to capture the correspondence between recipe components and the image at Title, Ingredient, and Instruction levels by measuring the mutual information. Extensive experiments demonstrate the effectiveness of UTE-FCL compared to existing methods.
Haruya Kyutoku, Keisuke Doman, Takahiro Komamizu, Ichiro Ide, Jiangbo Qian
Knowl. Based Syst.5
2024 Computational measurement of perceived pointiness from pronunciation
abstract
Abstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.4
2024 Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.4
2023 RecipeMeta: Metapath-enhanced Recipe Recommendation on Heterogeneous Recipe Network
abstract
Recipe is a set of instructions that describes how to make food. It can help people from the preparation of ingredients, food cooking process, etc. to prepare the food, and increasingly in demand on the Web. To help users find the vast amount of recipes on the Web, we address the task of recipe recommendation. Due to multiple data types and relationships in a recipe, we can treat it as a heterogeneous network to describe its information more accurately. To effectively utilize the heterogeneous network, metapath was proposed to describe the higher-level semantic information between two entities by defining a compound path from peer entities. Therefore, we propose a metapath-enhanced recipe recommendation framework, RecipeMeta, that combines GNN (Graph Neural Network)-based representation learning and specific metapath-based information in a recipe to predict User-Recipe pairs for recommendation. Through extensive experiments, we demonstrate that the proposed model, RecipeMeta, outperforms state-of-the-art methods for recipe recommendation.
Jialiang Shi, Takahiro Komamizu, Keisuke Doman, Haruya Kyutoku, Ichiro Ide
MMAsia5
2023 Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (1)5
2023 Discovering Phonesthemic Clusters in Readings of Kanji Characters toward Exploring Phonestheme in Japanese
Akira Yoshida, Chihaya Matsuhira, Hirotaka Kato, Takatsugu Hirayama, Takahiro Komamizu, Ichiro Ide
PACLIC6
2022 Detection of Birds in a 3D Environment Referring to Audio-Visual Information
abstract
We propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed.
Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi
AVSS2
2022 A Novel Approach for Pill-Prescription Matching with GNN Assistance and Contrastive Learning
Trung Thanh Nguyen 0006, Hoang Dang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Ichiro Ide, Phi-Le Nguyen
PRICAI (1)5
2021 MMArt-ACM'21: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2021
abstract
The International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and social media. The topics of the accepted papers cover an analytic topic on comic contents understanding to generative topics on image synthesis and conversion. The actual MMArt-ACM'21 Proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3460426.
Min-Chun Hu 0001, Ichiro Ide, Kensuke Tobitani
ICMR2
2021 Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
MMM (2)3
2021 Soft-Boundary Label Relaxation with class placement constraints for semantic segmentation of the railway environment
Yuki Furitsu, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Hiroki Mukojima, Nozomi Nagamine
Pattern Recognit. Lett.4
2020 LFIR2Pose: Pose Estimation from an Extremely Low-resolution FIR image Sequence
abstract
In this paper, we propose a method for human pose estimation from a Low-resolution Far-InfraRed (LFIR) image sequence captured by a 16 × 16 FIR sensor array. Human body estimation from such a single LFIR image is a hard task. For training the estimation model, annotation of the human pose to the images is also a difficult task for human. Thus, we propose the LFIR2Pose model which accepts a sequence of LFIR images and outputs the human pose of the last frame, and also propose an automatic annotation system for the model training. Additionally, considering that the scale of human body motion is largely different among body parts, we also propose a loss function focusing on the difference. Through an experiment, we evaluated the human pose estimation accuracy with an original data set, and confirmed that human pose can be estimated accurately from an LFIR image sequence.
Saki Iwata, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa
ICPR4
2020 Ω-GAN: Object Manifold Embedding GAN for Image Generation by Disentangling Parameters into Pose and Shape Manifolds
abstract
In this paper, we propose Object Manifold Embedding GAN (Ω-GAN) to generate images of variously shaped and arbitrarily posed objects from a noise variable sampled from a distribution defined over the pose and the shape manifolds in a vector space. We introduce Parametric Manifold Sampling to sample noise variables from a distribution over the pose manifold to conditionally generate object images in arbitrary poses by tuning the pose parameter. We also introduce Object Identity Loss for clearly disentangling the pose and shape parameters, which allows us to maintain the shape of the object instance when only the pose parameter is changed. Through evaluation, we confirmed that the proposed Ω-GAN could generate variously shaped object images in arbitrary poses by changing the pose and shape parameters independently. We also introduce an application of the proposed method for object pose estimation, through which we confirmed that the object poses in the generated images are accurate.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICPR3
2020 Median-Shape Representation Learning for Category-Level Object Pose Estimation in Cluttered Environments
abstract
In this paper, we propose an occlusion-robust pose estimation method of an unknown object instance in an object category from a depth image. In a cluttered environment, objects are often occluded mutually. For estimating the pose of an object in such a situation, a method that de-occludes the unobservable area of the object would be effective. However, there are two difficulties; occlusion causes the offset between the center of the actual object and its observable area, and different instances in a category may have different shapes. To cope with these difficulties, we propose a two-stage Encoder-Decoder model to extract features with objects whose centers are aligned to the image center. In the model, we also propose the Median-shape Reconstructor as the second stage to absorb shape variations in a category. By evaluating the method with both a large-scale virtual dataset and a real dataset, we confirmed the proposed method achieves good performance on pose estimation of an occluded object from a depth image.
Hiroki Tatemichi, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Ayako Amma, Hiroshi Murase
ICPR4
2020 MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020
abstract
The International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173
Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki
ICMR2
2020 CEA'20: The 12th Workshop on Multimedia for Cooking and Eating Activities
abstract
The 12th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'20 workshop and the list of papers presented in the workshop.
Ichiro Ide, Yoko Yamakata, Atsushi Hashimoto 0001
ICMR1
2020 Imageability Estimation using Visual and Language Features
abstract
Imageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability.
Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
ICMR3
2020 Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
MMM (2)2
2020 More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
MMM (2)3
2020 Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.2
2019 Exemplar-Based Pseudo-Viewpoint Rotation for White-Cane User Recognition from a 2D Human Pose Sequence
abstract
In recent years, various facilities are equipped to support visually impaired people, but accidents caused by visual disabilities still occur. In this paper, to support the visually-impaired people in a public space, we aim to classify whether a pedestrian image sequence obtained by a surveillance camera is a white-cane user or not from the temporal transition of a human pose represented as 2D coordinates. However, since the appearance of the 2D pose varies largely depending on the viewpoint of the pose, it is difficult to classify them. So, in this paper, we propose a method to rotate the viewpoint of a pose from various pseudo-viewpoints based on a pair of 2D poses simultaneously observed and classify the sequence by multiple classifiers corresponding to each viewpoint. Viewpoint rotation makes it possible to obtain pseudo-poses seen from various pseudo-viewpoints, extract richer pose features, and recognize white-cane users more accurately. Through an experiment, we confirmed that the proposed method improves the recognition rate by 12% compared to the method not employing viewpoint rotation.
Naoki Nishida 0003, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Jun Piao
AVSS4
2019 Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.2
2018 Gaze-Inspired Learning for Estimating the Attractiveness of a Food Photo
abstract
The number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them.
Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM5
2018 Voting-based Hand-Waving Gesture Spotting from a Low-Resolution Far-Infrared Image Sequence
abstract
We propose a temporal spotting method of a hand gesture from a low-resolution far-infrared image sequence captured by a far-infrared sensor array. The sensor array captures the spatial distribution of far-infrared intensity as a thermal image by detecting far-infrared waves emitted from heat sources. It is difficult to spot a hand gesture from a sequence of thermal images captured by the sensor due to its low-resolution, heavy noise, and varying duration of the gesture. Therefore, we introduce a voting-based approach to spot the gesture with template matching-based gesture recognition. We confirm the effectiveness of the proposed temporal spotting method in several settings.
Yasutomo Kawanishi, Chisato Toriyama, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa, Masato Kawade
VCIP5
2017 Action recognition from extremely low-resolution thermal image sequence
abstract
This paper proposes a Deep Learning-based action recognition method from an extremely low-resolution thermal image sequence. The method recognizes daily actions by humans (e.g. walking, sitting down, standing up, etc.) and abnormal actions (e.g. falling down) without privacy concerns. While privacy concerns can be ignored, it is difficult to compute feature points and to obtain a clear edge of the human body from an extremely low-resolution thermal image. To address these problems, this paper proposes a Deep Learning-based action recognition method that combines convolution layers and an LSTM layer for learning spatio-temporal representation, whose inputs are the thermal images and their frame differences cropped by the gravity center of human regions. The effectiveness of the proposed method was confirmed through experiments.
Takayuki Kawashima, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Daisuke Deguchi, Tomoyoshi Aizawa, Masato Kawade
AVSS3
2017 Automatic Selection of Web Contents Towards Automatic Authoring of a Video Biography
abstract
In this paper, we propose a method for image selection using Web image search for automatic video biography authoring. In the proposed method, images are selected from the image search results considering their visual contents for inclusion in the video biography. Through evaluation, we confirmed the effectiveness of the proposed image selection method compared to a baseline method which simply selects the top 1 search result.
Ichiro Ide, Yasutomo Kawanishi, Kyoka Kunishiro, Frank Nack, Daisuke Deguchi, Hiroshi Murase
ISM1
2017 Summarization of News Videos Considering the Consistency of Auditory and Visual Contents
abstract
Since news videos are valuable sources of multimedia information on real-world events, there is a demand for viewing them efficiently. However, there is a problem that summarization methods based on auditory contents do not take into account the visual contents. In the case of news videos, due to its presentation style where audio contents and visual contents do not necessarily come from the same source, this could severely decrease the amount of informative visual contents included in the generated summarized video. Thus, we propose a method for summarizing a sequence of news videos considering the consistency of both auditory and visual contents. The proposed method first selects key-sentences from the auditory contents (Closed Caption) of each news story in the sequence, and then selects a shot within the news story whose "Visual Concepts" detected from the visual contents are the most consistent with the key-phrase. Finally, the audio segment corresponding to each key-phrase is overlapped onto the selected shot, and then concatenated to generate a summarized video. The effectiveness of the proposed method was confirmed on several news topics through a subjective experiment.
Ichiro Ide, Ryunosuke Tanishige, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
ISM1
2017 Monocular localization within sparse voxel maps
abstract
We introduce a method that uses a single camera to localize a vehicle within a pre-constructed map consisting of a voxel occupancy grid and road-line marker positions. Sophisticated mapping hardware is capable of creating high-accuracy 3D maps of road environments, but localizing a vehicle within such maps is one of the challenges at the forefront of automated driving. A solution which is robust to dynamic environments, while using only inexpensive sensors, is a difficult problem. In addition, maps that enable precise localization consume a lot of data which is impractical for the expansive environments encountered in real-world road networks. We show how using the area of edge regions shared between rendered views of a compact voxel map and in-vehicle camera images can be coupled with non-linear optimization methods to determine the camera position and pose.
David Wong 0002, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium4
2017 Proposal of a spectral random dots marker using local feature for posture estimation
abstract
We propose a novel marker for robot's grasping task which has the following three aspects: (i) it is easy-to-find in a cluttered background, (ii) it is calculable for its posture (iii) its size is compact. The proposed marker is composed of a random dots pattern, and uses keypoint detection and a scale estimation by Spectral SIFT for dots detection and data decoding. The data is encoded by the scale size of dots, and the same dots in the marker work for both marker detection and data decoding. As a result, the proposed marker size can be compact. We confirmed the effectiveness of the proposed marker through experiments.
Norimasa Kobori, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
VR3
2017 Regression of feature scale tracklets for decimeter visual localization
David Wong 0002, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase
Image Vis. Comput.4
2016 A classification method of cooking operations based on eye movement patterns
abstract
We are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables.
Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ETRA5
2016 Moving camera background-subtraction for obstacle detection on railway tracks
abstract
We propose a method for detecting obstacles by comparing input and reference train frontal view camera images. In the field of obstacle detection, most methods employ a machine learning approach, so they can only detect pre-trained classes, such as pedestrian, bicycle, etc. This means that obstacles of unknown classes cannot be detected. To overcome this problem, we propose a background subtraction method that can be applied to moving cameras. First, the proposed method computes frame-by-frame correspondences between the current and the reference (database) image sequences. Then, obstacles are detected by applying image subtraction to corresponding frames. To confirm the effectiveness of the proposed method, we conducted an experiment using several image sequences captured on an experimental track. Its results showed that the proposed method could detect various obstacles accurately and effectively.
Hiroki Mukojima, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Masato Ukai, Nozomi Nagamine, Ryuta Nakasone
ICIP4
2016 Misclassification tolerable learning for robust pedestrian orientation classification
abstract
In this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
ICPR3
2016 Parts Selective DPM for detection of pedestrians possessing an umbrella
abstract
In recent years, pedestrian detection from an in-vehicle camera has been attracting attention. However, in the case of a raining situation, the detection accuracy decreases because the head of a pedestrian tends to be occluded by an umbrella. In oder to handle such cases, in this paper, as a variation of the Deformable Part Model (DPM) which is widely used in the field of object recognition, we propose “Parts Selective DPM (PS-DPM)” which selectively chooses the original part filters and additional part filters trained independently. In the detection of pedestrians possessing an umbrella, the selection of head and umbrella parts will make pedestrian detection more robust to the occlusion. We conducted experiments to evaluate the performance of the proposed method. As a result, pedestrian detection with the proposed PS-DPM achieved high detection accuracy in rainy weather, compared with the detection by the conventional DPM. Moreover, we confirmed that it did not decrease the pedestrian detection accuracy in fine weather.
Yuto Shimbo, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium4
2015 Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF camera
abstract
This paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification.
Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
Intelligent Vehicles Symposium4
2014 Spatial People Density Estimation from Multiple Viewpoints by Memory Based Regression
abstract
Crowd analysis using cameras has attracted much attention for public safety and marketing. Among techniques of the crowd analysis, we focus on spatial people density estimation which estimates the number of people for each small area in a floor region. However, spatial people density cannot be estimated accurately for an area far from the camera because of the occlusion by people in a closer area. Therefore, we propose a method using a memory based regression method with images captured from cameras from multiple viewpoints. This method is realized by looking up a table that consists of correspondences between people density maps and crowd appearances. Since the crowd appearances include situations where various occlusions occur, an estimation robust to occlusion should be realized. In an experiment, we examined the effectiveness of the proposed method.
Yoshimune Tabuchi, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takayuki Kurozumi, Kunio Kashino
ICPR4
2014 Scene Duplicate Detection from News Videos Using Image-Audio Matching Focusing on Human Faces
abstract
As one tool for structuring a massive volume of archived news videos based on their semantic contents, this paper proposes a method to detect scene duplicates from news videos. A scene duplicate is a pair of video segments taken at the same event from different viewpoints. Referring to the audio channel is effective to detect scene duplicates regardless of viewpoints, but it cannot be relied on when external audio sources (e.g. Narrations, sound effects) overlap the original one. In contrast, the image channel can be useful in most cases, although significant difference in viewpoints affect the detection. The proposed method integrates the information from these two channels in order to improve the accuracy of scene duplicate detection from news videos. The performance of the proposed method was evaluated through an experiment with actual broadcast news videos. As a result, we obtained the higher detection accuracies in both recall and precision. Therefore, we confirmed the effectiveness of the proposed method.
Haruka Kumagai, Keisuke Doman, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM3
2014 Estimation of traffic sign visibility considering local and global features in a driving environment
abstract
This paper proposes a camera-based visibility estimation method for a traffic sign. The visibility here indicates how a visual target is easy to be detected and recognized by a human driver (not a machine). This research aims at realizing a nuisance-free driver assistance system which sorts out information depending on the visibility of a visual target, in order to prevent driver distraction. Our previous study on estimating the visibility of a traffic sign considered only the effect of the local region around a target, assuming the situation that a driver's gaze is around it. The proposed method integrates both the local features and global features in a driving environment without such an assumption. The global features evaluate the positional relationships between traffic signs and the appearance around the fixation point of a driver's gaze, which considers the effect of the driver's entire field of view. Experimental results showed the effectiveness of incorporating the global features for estimating the visibility of a traffic sign.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Utsushi Sakai
Intelligent Vehicles Symposium5
2014 Single camera vehicle localization using SURF scale and dynamic time warping
abstract
Vehicle ego-localization is an essential process for many driver assistance and autonomous driving systems. The traditional solution of GPS localization is often unreliable in urban environments where tall buildings can cause shadowing of the satellite signal and multipath propagation. Typical visual feature based localization methods rely on calculation of the fundamental matrix which can be unstable when the baseline is small. In this paper we propose a novel method which uses the scale of matched SURF image features and Dynamic Time Warping to perform stable localization. By comparing SURF feature scales between input images and a pre-constructed database, stable localization is achieved without the need to calculate the fundamental matrix. In addition, 3D information is added to the database feature points in order to perform lateral localization, and therefore lane recognition. From experimental data captured from real traffic environments, we show how the proposed system can provide high localization accuracy relative to an image database, and can also perform lateral localization to recognize the vehicle's current lane.
David Wong 0002, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium3
2014 Estimation of the Representative Story Transition in a Chronological Semantic Structure of News Topics
abstract
It is important to track the flow of topics to thoroughly understand the contents. Accordingly, a method that structures the chronological semantic relations between news stories, namely a "topic thread structure" has been proposed. It allows the comprehensive understanding of a topic by chronologically tracking stories one by one from the initial story. However, this task imposes a user to watch many stories when it contains various sub-topics. Thus, we propose a method that estimates the representative story transition in a topic thread structure. In the proposed method, features obtained from a story and those from the topic thread structure are used for the estimation. We confirmed the effectiveness of the proposed method by comparing the results obtained from the proposed method to the ground truth obtained from votes in a subjective experiment.
Kosuke Kato, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ICMR2
2014 Event Detection based on Twitter Enthusiasm Degree for Generating a Sports Highlight Video
abstract
This paper presents a Twitter-based event detection method based on "Twitter Enthusiasm Degrees (TED)" toward generating a highlight video of a sports game. Existing methods not only depend on both languages and sports types but also often falsely detect non-target events. In contrast, the proposed method detects sports events using TEDs calculated from several kinds of string features independent of languages and sports. We applied the proposed method to actual sports games, and compared the detected events with the events present in broadcasted highlight videos, and confirmed the effectiveness and the language and sports type independencies of the proposed method.
Keisuke Doman, Taishi Tomita, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ACM Multimedia3
2013 Pedestrian detection by scene dependent classifiers with generative learning
abstract
Recently, pedestrian detection from in-vehicle camera images is becoming an crucial technology for Intelligent Transportation Systems (ITS). However, it is difficult to detect pedestrians accurately in various scenes by obtaining training samples. To tackle this problem, we propose a method to construct scene dependent classifiers to improve the accuracy of pedestrian detection. The proposed method selects an appropriate classifier based on the scene information that is a category of appearance associated with location information. To construct scene dependent classifiers, the proposed method introduces generative learning for synthesizing scene dependent training samples. Experimental results showed that the detection accuracy of the proposed method outperformed the comparative method, and we confirmed that scene dependent classifiers improved the accuracy of pedestrian detection.
Hidefumi Yoshida, Daichi Suzuo, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takashi Machida, Yoshiko Kojima
Intelligent Vehicles Symposium4
2013 Recompilation of Broadcast Videos Based on Real-World Scenarios
Ichiro Ide
MMM (2)1
2013 Detection of Biased Broadcast Sports Video Highlights by Attribute-Based Tweets Analysis
Takashi Kobayashi 0001, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
MMM (2)4
2012 Robust Face Super-Resolution Using Free-Form Deformations for Low-Quality Surveillance Video
abstract
Recently, the demand for face recognition to identify persons from surveillance video cameras has rapidly increased. Since surveillance cameras are usually placed at positions far from a person's face, the quality of face images captured by the cameras tends to be low. This degrades the recognition accuracy. Therefore, aiming to improve the accuracy of the low-resolution-face recognition, we propose a video-based super-resolution method. The proposed method can generate a high-resolution face image from low-resolution video frames including non-rigid deformations caused by changes of face poses and expressions without using any positional information of facial feature points. Most existing techniques use the facial feature points for image alignment between the video frames. However, it is difficult to obtain the accurate positions of the feature points from low-resolution face images. To achieve the alignment, the proposed method uses a free-form deformation method that flexibly aligns each local region between the images. This enables super-resolution of face images from low-resolution videos. Experimental results demonstrated that the proposed method improved the performance of super-resolution for actual videos in terms of both image quality and face recognition accuracy.
Tomonari Yoshida, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICME4
2012 Estimation of the human performance for pedestrian detectability based on visual search and motion features
Masashi Wakayama, Daisuke Deguchi, Keisuke Doman, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
ICPR4
2012 Smart VideoCooKing: a multimedia cooking recipe browsing application on portable devices
abstract
This demo presents "Smart VideoCooKing" which is a multimedia cooking recipe browsing application on portable Android devices. A multimedia cooking recipe is a cooking recipe where each cooking operation is associated with a corresponding video clip describing it, aimed to facilitate the understanding of cooking operations. In combination with third-party applications, "Smart VideoCooKing" provides useful functions such as playing cooking video clips describing cooking operations quickly, searching information of ingredients easily, and reading aloud cooking directions.
Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ACM Multimedia4
2012 Overview of the ACM multimedia 2012 workshop on multimedia for cooking and eating activities (CEA'12)
abstract
This overview introduces the aim of the CEA'12 workshop and the list of papers presented in the workshop.
Mutsuo Sano, Ichiro Ide, Yoko Yamakata
ACM Multimedia2
2011 Low Resolution QR-Code Recognition by Applying Super-Resolution Using the Property of QR-Codes
abstract
This paper proposes a method for low resolution QR-code recognition. A QR-code is a two-dimensional binary symbol that can embed various information such as characters and numbers. To recognize a QR-code correctly and stably, the resolution of an input image should be high. In practice, however, recognition of a QR-code is usually difficult due to low resolution when it is captured from a distance. In this paper, we propose a method to improve the performance of low resolution QR-code recognition by using the super-resolution technique that generates a high resolution image from multiple low-resolution images. Although a QR-code is a binary pattern, it is observed as a grayscale image due to the degradation through the capturing process. Especially the pixels around the borders between white and black regions become ambiguous. To overcome this problem, the proposed method introduces a binary pattern constraint to generate super-resolved images appropriate for recognition. Experimental results showed that a recognition rate of 98% can be achieved by the proposed method, which is a 15.7% improvement in comparison with a method using a conventional super-resolution method.
Yuji Kato, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR4
2011 Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News Videos
abstract
We propose a method to detect the inconsistency between a subject and the speaker for extracting speech scenes from news videos. Speech scenes in news videos contain a wealth of multimedia information, and are valuable as archived material. In order to extract speech scenes from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such approach, since news videos contain non-speech scenes where the speaker is not the subject, such as narrated scenes. To solve this problem, we propose a method to discriminate between speech scenes and narrated scenes based on the co-occurrence between a subject's lip motion and the speaker's voice. The proposed method uses lip shape and degree of lip opening as visual features representing a subject's lip motion, and uses voice volume and phoneme as audio feature representing a speaker's voice. Then, the proposed method discriminates between speech scenes and narrated scenes based on the correlations of these features. We report the results of experiments on videos captured in a laboratory condition and also on actual broadcast news videos. Their results showed the effectiveness of our method and the feasibility of our research goal.
Shogo Kumagai, Keisuke Doman, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM5
2011 Intelligent traffic sign detector: Adaptive learning based on online gathering of training samples
abstract
This paper proposes an intelligent traffic sign detector using adaptive learning based on online gathering of training samples from in-vehicle camera image sequences. To detect traffic signs accurately from in-vehicle camera images, various training samples of traffic signs are needed. In addition, to reduce false alarms, various background images should also be prepared before constructing the detector. However, since their appearances vary widely, it is difficult to obtain them exhaustively by manual intervention. Therefore, the proposed method simultaneously obtains both traffic sign images and background images from in-vehicle camera images. Especially, to reduce false alarms, the proposed method gathers background images that were easily mis-detected by a previously constructed traffic sign detector, and re-trains the detector by using them as negative samples. By using retrospectively tracked traffic sign images and background images as positive and negative training samples, respectively, the proposed method constructs a highly accurate traffic sign detector automatically. Experimental results showed the effectiveness of the proposed method.
Daisuke Deguchi, Daisuke Shirasuna, Keisuke Doman, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium4
2011 Estimation of traffic sign visibility considering temporal environmental changes for smart driver assistance
abstract
We propose a visibility estimation method for traffic signs considering temporal environmental changes, as a part of work for the realization of nuisance-free driver assistance systems. Recently, the number of driver assistance systems in a vehicle is increasing. Accordingly, it is becoming important to sort out appropriate information provided from them, because providing too much information may cause driver distraction. To solve such a problem, we focus on a visibility estimation method for controlling the information according to the visibility of a traffic sign. The proposed method sequentially captures a traffic sign by an in-vehicle camera, and estimates its accumulative visibility by integrating a series of instantaneous visibility. By this way, even if the environmental conditions may change temporally and complicatedly, we can still accurately estimate the visibility that the driver perceives in an actual traffic scene. We also investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium5
2011 Road image update using in-vehicle camera images and aerial image
abstract
Road image is becoming important for several applications such as car navigation systems, traffic environment research, city modeling. Usually, a road image can be obtained from an aerial image but the resolution of the aerial image is often low, or it contains occlusions by obstacles. Therefore, the update of road image is required. In this paper, we propose a road image mosaicing method using in-vehicle camera images and an aerial image. We first perform image registration of road regions between these images, and then, we generate a large road image by performing image mosaicing of road regions in invehicle camera images. In an experiment, we achieved resolution improvement and occlusions removal, and also succeeded in update of a large road image.
Masafumi Noda, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Yoshiko Kojima, Takashi Naito
Intelligent Vehicles Symposium4
2011 3-D line segment reconstruction using an in-vehicle camera for free space detection
abstract
Free space detection is very important for vehicle navigation and safe driving. 3-D line segment reconstruction of a street is important for the free space detection because a street-view includes many line segments. For the free space detection, we propose a method for reconstructing 3-D line segments in a streetscape using a monocular in-vehicle camera. The 3-D reconstruction of the line segments is achieved by using each three images from an image sequence. Once accurate camera poses of these images are obtained, one of the remaining crucial problems is to match the line segments between the images correctly. A strategy for finding correspondence of the line segments is as follows: First, the correspondences of line segment candidates are searched by using a two-view constraint. However, the two-view constraint has difficulty on determining an unique correspondence geometrically. Therefore, the candidates of the line segment correspondences are reduced using a three-view constraint. In order to improve the accuracy, the proposed method exploits a color feature of the line segment and a preliminary knowledge of the vehicle motion. Finally, the line segments are reconstructed using the correspondences. From an experimental result, we confirmed the effectiveness of the proposed method. Application to the free space detection demonstrated the usefulness of the reconstructed line segments.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium4
2011 Why did the prime minister resign?: generation of event explanations from large news repositories
abstract
One of the common parts of news is to provide the background for a current event, such as the resignation of a Prime Minister. This paper addresses a framework that facilitates semi-automated authoring of explanatory audio-visual news topics in a retrospective style for the domain of politics based on already edited new stories available in the repository of the news corporation. The aim is to facilitate a journalist with an audio-visual body based on which he/she can finalize the explanatory piece. The proposed framework enhances current state of the art video summarization by allowing the combination of different news stories into one coherent explanation about a topic of the current news. The framework introduces techniques that exploit demoscopic data in form of polls for the development of the general story outline; the automatic retrieval of relevant material by using a combination of event templates and automatic news summarization over topic threads; and the generation of the final video by applying a set of trimming rules. Example generations are presented and discussed and an outline of future work is presented.
Frank Nack, Ichiro Ide
ACM Multimedia2
2011 Scene segmentation of wedding party videos by scenario-based matching with example videos
abstract
We propose a method for scene segmentation of a wedding party video. Recently, it has become popular to take videos of a wedding ceremony and its party. Especially, because of its length, each scene of a wedding party video needs to be indexed with each event for efficient browsing. The proposed method segments a wedding party video into scenes of events by scenario- based matching with example videos that are synthesized by combining scenes from other wedding party videos according to a scenario.
Kazuki Sawai, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ACM Multimedia4
2011 Video CooKing: Towards the Synthesis of Multimedia Cooking Recipes
Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM (2)4
2010 Classification of Near-Duplicate Video Segments Based on Their Appearance Patterns
abstract
We propose a method that analyzes the structure of a large volume of general broadcast video data by the appearance patterns of near-duplicate video segments. We define six classification rules based on the appearance patterns of near-duplicate video segments according to their roles, and evaluated them over more than 1,000 hours of actual broadcast video data.
Ichiro Ide, Yuji Shamoto, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ICPR1
2010 Region-Based Image Transform for Transition Between Object Appearances
abstract
We propose a method of region-based image transform to achieve accurate transition between object appearances. A view-transition model (VTM) is one of the statistical methods that learn appearance transition from a sample image dataset of a large number of objects with various appearances. However, the VTM method has a practical problem that the appearance transition cannot be performed accurately if a sufficient number of learning samples is not available in the dataset. To cope with the problem, the proposed method first determines the regions of input and output images whose pixel values mutually affect each other during appearance transition, then transforms iteratively between partial images in the regions. We conducted experiments using actual image datasets. The results show that the proposed method could accurately transform appearances compared with the VTM method.
Tomokazu Takahashi, Yuki Kono, Ichiro Ide, Hiroshi Murase
ICPR3
2010 Removal of Moving Objects from a Street-View Image by Fusing Multiple Image Sequences
abstract
We propose a method to remove moving objects from an in-vehicle camera image sequence by fusing multiple image sequences. Driver assistance systems and services such as Google Street View require images containing no moving object. The proposed scheme consists of three parts: (i) collection of many image sequences along the same route by using vehicles equipped with an omni-directional camera, (ii) temporal and spatial registration of image sequences, and (iii) mosaicing partial images containing no moving object. Experimental results show that 97.3% of the moving object area could be removed by the proposed method.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICPR4
2010 Multimedia Supplementation to a Cooking Recipe Text for Facilitating Its Understanding to Inexperienced Users
abstract
Assisting culinary activities for inexperienced users has been considered as an important task in most existing works in the field. On the other hand, recipe texts are becoming available on the Internet in increasing numbers. However, they tend to be written simply by mostly non-professional people, and thus are sometimes difficult for an inexperienced person to follow the steps and manage to cook as they are supposed to. In this paper, we propose a method that detects difficult descriptions for an inexperienced user in an existing text recipe, and supplements them with multimedia contents including text information extracted from a large number of recipes, and also images and video clips on certain kinds of cooking operations, to facilitate the understanding of the recipe. Experimental results showed promising ability of the proposed method to assist inexperienced users understand the descriptions in a recipe.
Ichiro Ide, Yuka Shidochi, Yuichi Nakamura 0001, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ISM1
2010 Estimation of traffic sign visibility toward smart driver assistance
abstract
We propose a visibility estimation method for traffic signs as part of work for realization of nuisance-free driving safety support systems. Recently, the number of driving safety support systems in a car has been increasing. As a result, it is becoming important to select appropriate information from them for safe and comfortable driving because too much information may cause driver distraction and may increase the risk of a traffic accident. One of the approaches to avoid such a problem is to alert the driver only with information which could easily be missed. Therefore, to realize such a system, we focus on estimating the visibility of traffic signs. The proposed method is a model-based method that estimates the visibility of traffic signs focusing on the difference of image features between a traffic sign and its surrounding region. In this paper, we investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium5
2010 PageRank with Text Similarity and Video Near-Duplicate Constraints for News Story Re-ranking
Xiaomeng Wu, Ichiro Ide, Shin'ichi Satoh 0001
MMM2
2010 A Hilbert warping method for handwriting gesture recognition
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
Pattern Recognit.3
2009 Low-Resolution Character Recognition by Video-Based Super-Resolution
abstract
In this paper, we propose a method for recognizing low-resolution characters using a super-resolution technique. Although portable digital cameras can be used for camera based character recognition, the captured images contain several types of noises which make the recognition task difficult. We introduce a phase of super-resolution before the recognition to enhance the resolution of images obtained from a video. The proposed method uses the subspace method for the recognition of characters which are integrated from multiple low-resolution characters by the super-resolution technique. Experimental results show that the proposed method improves the recognition accuracy; we confirmed that the recognition rate for the input size of 7 times 7 pixels was 90.35%, and for the input size of 9 times 9 pixels was 99.97%.
Ataru Ohkura, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR4
2009 Adaptive division of feature space for rapid detection of near-duplicate video segments
abstract
Near-duplicate video detection is becoming a core-technology for analyzing the structure of a large-scale video archive. It, however, is naturally an O(n2) problem, where n is a value proportional to the total length of an input video stream. We have previously challenged this time-consuming task by reducing the cost required for each of the O(n2) comparisons. This paper, on the other hand, proposes a method that reduces the number of comparisons by adaptively dividing the feature space according to the distribution of feature points.
Ichiro Ide, Shugo Suzuki, Tomokazu Takahashi, Hiroshi Murase
ICME1
2009 Labeling News Topic Threads with Wikipedia Entries
abstract
Wikipedia is a famous online encyclopedia. However most Wikipedia entries are mainly explained by text, so it will be very informative to enhance the contents with multimedia information such as videos. Thus we are working on a method to extend information of Wikipedia entries by means of broadcast videos which explain the entries. In this work, we focus especially on news videos and Wikipedia entries about news events. In order to extend information of Wikipedia entries, it is necessary to link news videos and Wikipedia entries. So the main issue will be on a method that labels news videos with Wikipedia entries automatically. In this way, explanations could be more detailed with news videos can be exhibited, and the context of the news events should become easier to understand. Through experiments, news videos were accurately labeled with Wikipedia entries with a precision of 86% and a recall of 79%.
Tomoki Okuoka, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM4
2009 Overview of the ACM multimedia 2009 workshop on multimedia for cooking and eating activities (CEA'09)
abstract
This overview introduces the aim of the CEA'09 workshop and the list of papers presented in the workshop.
Mutsuo Sano, Ichiro Ide, Kenzaburo Miyawaki
ACM Multimedia2
2009 A Multimodal Constellation Model for Object Category Recognition
Yasunori Kamiya, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM3
2008 A Hilbert Warping Algorithm for Recognizing Characters from Moving Camera
abstract
We present a method for recognizing characters from image sequences captured by moving camera. In the proposed method, the sequence of the captured images is compared with those of reference character patterns using the concept of analytic signal. Since the captured image sequence can be nonlinearly warped along the time axis due to the movement of a hand-held camera, phase synchronization of two analytic signals is used for the alignment of two image sequences. Hilbert transform is used to convert all the image sequences into analytic signals whose phases are supposed to be increasing. Experimental results showed the usefulness of the proposed phase-based alignment algorithm.
Hiroyuki Ishida, Ichiro Ide, Hiroshi Murase, Tomokazu Takahashi
Document Analysis Systems2
2008 A Hilbert warping method for camera-based finger-writing recognition
abstract
We propose a time-warping algorithm for recognizing finger actions by a camera. In the proposed method, an input image sequence is aligned to the reference sequences by phase-synchronization of the analytic signals, and then classified by comparing the cumulative distances. A major benefit of this method is that over-fitting to sequences of incorrect categories is restricted. The proposed method exhibited high recognition accuracy in finger-writing character recognition.
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICPR3
2008 Eigenspace interpolation for appearance-based object recognition
abstract
An eigenspace interpolation method smoothly interpolates between two different eigenspaces using high dimensional rotation. However, up to now its effectiveness in object recognition and the validity of the interpolation algorithm have not been discussed sufficiently. We therefore propose an appearance-based object recognition method combining the eigenspace interpolation method and a subspace method. We conducted face recognition experiments using images captured from multiple camera positions with various illumination conditions. Experimental results demonstrate the effectiveness of the proposed method and the validity of the interpolation algorithm.
Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
ICPR3
2008 Cross-Lingual Retrieval of Identical News Events by Near-Duplicate Video Segment Detection
Akira Ogawa, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM3
2008 Recognition of camera-captured low-quality characters using motion blur information
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
Pattern Recognit.3
2007 Interpolation Between Eigenspaces Using Rotation in Multiple Dimensions
Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
ACCV (2)3
2007 Genre-Adaptive Near-Duplicate Video Segment Detection
abstract
This paper proposes a fast and accurate method to detect all near-duplicate segments in a video stream. To reduce the computation time while ensuring the detection accuracy equivalent to that by brute-force frame-by-frame comparison, a two-step detection method is proposed; a fast but rough detection applied in a compressed feature vector space spanned by the result of a PC A, followed by confirmation of candidates in the original high dimension space. The results show that the proposed method accelerates the detection by more than 1,000 times while maintaining the detection accuracy. We also propose an entropy-based pixel selection scheme to generate feature vectors optimized for comparison of video segments within programs with mostly common pictures. The results show that the proposed scheme eliminates the false positives drastically, which should lead to even faster detection.
Ichiro Ide, Kazuhiro Noda, Tomokazu Takahashi, Hiroshi Murase
ICME1
2007 mediaWalker: a video archive explorer based on time-series semantic structure
abstract
We introduce a video browsing interface 'mediaWalker' that lets users explore a news video archive based on a time-series semantic structure; the 'topic thread' structure. The interface lets users efficiently track up and down the development of news in an archive with more than 1,000 hours of video.
Ichiro Ide, Tomoyoshi Kinoshita, Tomokazu Takahashi, Shin'ichi Satoh 0001, Hiroshi Murase
ACM Multimedia1
2006 Spatiotemporal Density Feature Analysis to Detect Liver Cancer from Abdominal CT Angiography
Yoshito Mekada, Yuki Wakida, Yuichiro Hayashi, Ichiro Ide, Hiroshi Murase
ACCV (2)4
2006 Exploiting Topic Thread Structures in a News Video Archive for the Semi-Automatic Generation of Video Summaries
abstract
We propose a method that semi-automatically composes video stories by connecting individual stories in a news video archive along a topic-based semantic structure, namely the topic thread. We introduce the methods to realize the composition, namely, story segmentation, topic threading and clustering. We then evaluate the proposed approach based on preliminary tests. Since the thread structure reflects the development of topics in the real-world, we believe that the composed news video story should be effective for the user to gain a deeper understanding of the current topic of interest
Ichiro Ide, Hiroshi Mo, Norio Katayama, Shin'ichi Satoh 0001
ICME1
2005 Automated Nomenclature of Bronchial Branches Extracted from CT Images and Its Application to Biopsy Path Planning in Virtual Bronchoscopy
Kensaku Mori, Sinya Ema, Takayuki Kitasaka, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori
MICCAI (2)5
2005 Cooking navi: assistant for daily cooking in kitchen
abstract
We are developing a cooking navigation system, which helps even a novice user to cook several recipes in parallel without failure, while improving an advanced user's skill further. To realize this, the system optimizes the cooking procedure considering the following restrictions: (1) Duration of cooking, (2) Accuracy of cooking, and (3) Learning effect, by providing appropriate instructions to user's at the right timing, making full use of multimedia information. The users should be able to cook perfectly and comfortably just by following the text, video and audio provided by the system. According to the result of a preliminary experiment, all users from novice to experienced cooks could finish two dishes in parallel while enjoyeing the cooking very much. The result of a questionnaire shows the effectiveness of the multimedia navigation that we propose.
Reiko Hamada, Jun Okabe, Ichiro Ide, Shin'ichi Satoh 0001, Shuichi Sakai, Hidehiko Tanaka
ACM Multimedia3
2003 Topic-based inter-video structuring of a large-scale news video corpus
abstract
We propose a topic-based inter-video news video corpus structuring method and a visual interface to efficiently browse through the structured corpus. Such inter-video structuring was not deeply sought in previous works. The topic-based structure is analyzed by closed-caption text analysis; topic segmentation and tracking. The visual interface provides the ability to 1) search and select a topic by query terms and 2) track a topic thread interactively referring to the text analysis results. Although topic retrieval is somewhat similar to conventional video retrieval methods, the combination with topic tracking makes it remarkably easy to narrow down the results that match a user's interest and moreover reveal underlying content-based structures, where the structure itself contains rich information.
Ichiro Ide, Hiroshi Mo, Norio Katayama, Shin'ichi Satoh 0001
ICME1
2002 An object detection method for describing soccer games from video
abstract
We propose a novel object detection and tracking method in order to detect and track objects necessary to describe contents of a soccer game. On the contrary to intensity oriented conventional object detection methods, the proposed method refers to color rarity and local edge property, and integrally evaluates them by a fuzzy function to achieve better detection quality. These image features were chosen considering the characteristics of soccer video images, that most non-object regions are roughly single colored (green) and most objects tend to have locally strong edges. We also propose a simple object tracking method, that could track objects with occlusion with other objects using a color based template matching. The result of an evaluation experiment applied to actual soccer video showed very high detection rate in detecting player regions without occlusion, and promising ability for regions with occlusion.
Okihisa Utsumi, Koichi Miura, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka
ICME (1)3
1999 Associating video with related documents
abstract
Article Free Access Share on Associating video with related documents Authors: Reiko Hamada Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Ichiro Ide Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Shuichi Sakai Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Hidehiko Tanaka Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999Pages 17–20https://doi.org/10.1145/319878.319883Published:01 October 1999Publication History 1citation199DownloadsMetricsTotal Citations1Total Downloads199Last 12 Months11Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Reiko Hamada, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka
ACM Multimedia (2)2