VLDB 2026 Research / reviewers in the wild / expert
Ichiro Ide
dblp:33/339
· DBLP profile ↗
95ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0003-3942-9296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 37 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | View-aware Cross-modal Distillation for Multi-view Action RecognitionabstractThe widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
WACV | 5 |
| 2026 | MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environmentsabstractMulti-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
Pattern Recognit. | 5 |
| 2026 | Hierarchical Global-Local Fusion for One-stage Open-vocabulary Temporal Action DetectionabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) extends the detection scope of Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) to unseen action classes specified by vocabularies not included in the training data, within untrimmed video. Typical Open-vocab TAD methods adopt a two-stage approach that first proposes candidate action intervals and then identifies those actions. However, errors in the first stage can affect the subsequent stage and the final detection results. Moreover, conventional methods for temporal context analyses tend to focus solely on either global or local context. Focusing solely on the global context can lead to lack of momentary detail, making it difficult to distinguish one action from another. Conversely, focusing only on the local context makes it challenging to determine the start and end timings of action intervals. To address these challenges, we introduce a one-stage approach named Hierarchical Open-vocab TAD (HOTAD), consisting of two branches: Temporal Context Analysis (TCA) and Video–Text Alignment (VTA). The former utilizes Hierarchical Encoder (HE) to fuse global and local temporal features, enabling a comprehensive capture of temporal actions, while the latter branch exploits the synergy between visual and textual modalities for precisely detecting unseen actions in the Open-vocab setting. Experiments and in-depth analysis using the widely recognized datasets THUMOS14 and ActivityNet-1.3 are performed to show the effectiveness of HOTAD. The results highlight remarkable accuracy in detecting a wide range of unseen actions. Furthermore, HOTAD significantly reduces wrong labels and localizes action instances with high precision, showcasing its robustness in complex and dynamic video settings. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor FusionabstractMulti-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
FG | 5 |
| 2025 | Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video CaptioningabstractRecent advances in video captioning are driven by large-scale pretrained models, which follow the standard “pre-training followed by fine-tuning” paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) approach offers a promising alternative, but primarily focuses on the language components of Multimodal Large Language Models (MLLMs). Despite recent progress, PEFT remains underexplored in multimodal tasks and lacks sufficient understanding of visual information during fine-tuning the model. To bridge this gap, we propose Query-Adapter (Q-Adapter), a lightweight visual adapter module designed to enhance MLLMs by enabling efficient fine-tuning for the video captioning task. Q-Adapter introduces learnable query tokens and a gating layer into Vision Encoder, enabling effective extraction of sparse, caption-relevant features without relying on external textual supervision. We evaluate Q-Adapter on two well-known video captioning datasets, MSR-VTT and MSVD, where it achieves state-of-the-art performance among the methods that take the PEFT approach across BLEU@4, METEOR, ROUGE-L, and CIDEr metrics. Q-Adapter also achieves competitive performance compared to methods that take the full fine-tuning approach while requiring only 1.4% of the parameters. We further analyze the impact of key hyperparameters and design choices on fine-tuning effectiveness, providing insights into optimization strategies for adapter-based learning. These results highlight the strong potential of Q-Adapter in balancing caption quality and parameter efficiency, demonstrating its scalability for video–language modeling. Junan Chen 0004, Trung Thanh Nguyen 0006, Takahiro Komamizu, Ichiro Ide |
MMAsia | 4 |
| 2025 | Quantifying Image-Adjective Associations by Leveraging Large-Scale Pretrained Models
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide |
MMM (4) | 5 |
| 2025 | Towards Visual Storytelling by Understanding Narrative Context Through Scene-Graphs
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMM (4) | 5 |
| 2024 | R-DiP: Re-ranking Based Diffusion Pre-computation for Image Retrieval
Tatsuya Kato, Takahiro Komamizu, Ichiro Ide |
DEXA (2) | 3 |
| 2024 | One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-Scale and Action Label FeaturesabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing and classifying actions based on a predefined set of categories. In contrast, Open-vocab TAD goes further and is not limited to these predefined categories. This is particularly useful in real-world scenarios where the variety of actions in videos can be vast and not always predictable. The prevalent methods in Open-vocab TAD typically employ a 2-stage approach, which involves generating action proposals and then identifying those actions. However, errors made during the first stage can adversely affect the subsequent action identification accuracy. Additionally, existing studies face challenges in handling actions of different durations owing to the use of fixed temporal processing methods. Therefore, we propose a L-stage approach consisting of two primary modules: Multi-scale Video Analysis (MVA) and Video-Text Alignment (VTA). The MVA module captures actions at varying temporal resolutions, overcoming the challenge of detecting actions with diverse durations. The VTA module leverages the synergy between visual and textual modalities to precisely align video segments with corresponding action labels, a critical step for accurate action identification in Open-vocab scenarios. Evaluations on widely recognized datasets THUMOSl4 and ActivityNet-I.3, showed that the proposed method achieved superior results compared to the other methods in both Open-vocab and Closed-vocab settings. This serves as a strong demonstration of the effectiveness of the proposed method in the TAD task. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
FG | 4 |
| 2024 | Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image GenerationabstractText-to-image diffusion models sometimes depict blended concepts in the generated images. One promising use case of this effect would be the nonword-to-image generation task which attempts to generate images intuitively imaginable from a non-existing word (nonword). To realize nonword-to-image generation, an existing study focused on associating nonwords with similar-sounding words. Since each nonword can have multiple similar-sounding words, generating images containing their blended concepts would increase intuitiveness, facilitating creative activities and promoting computational psycholinguistics. Nevertheless, no existing study has quantitatively evaluated this effect in either diffusion models or the nonword-to-image generation paradigm. Therefore, this paper first analyzes the conceptual blending in a pretrained diffusion model, Stable Diffusion. The analysis reveals that a high percentage of generated images depict blended concepts when inputting an embedding interpolating between the text embeddings of two text prompts referring to different concepts. Next, this paper explores the best text embedding space conversion method of an existing nonword-to-image generation framework to ensure both the occurrence of conceptual blending and image generation quality. We compare the conventional direct prediction approach with the proposed method that combines k-nearest neighbor search and linear regression. Evaluation reveals that the enhanced accuracy of the embedding space conversion by the proposed method improves the image generation quality, while the emergence of conceptual blending could be attributed mainly to the specific dimensions of the high-dimensional text embedding space. Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide |
ACM Multimedia | 5 |
| 2024 | Action Selection Learning for Multi-label Multi-view Action RecognitionabstractMulti-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMAsia | 4 |
| 2024 | Cross-modal recipe retrieval based on unified text encoder with fine-grained contrastive learningabstractCross-modal recipe retrieval is vital for transforming visual food cues into actionable cooking guidance, making culinary creativity more accessible. Existing methods separately encode the recipe Title, Ingredient, and Instruction using different text encoders, then aggregate them to obtain recipe feature, and finally match it with encoded image feature in a joint embedding space. These methods perform well but require significant computational cost. In addition, they only consider matching the entire recipe and the image but ignore the fine-grained correspondence between recipe components and the image, resulting in insufficient cross-modal interaction. To this end, we propose U nified T ext E ncoder with F ine-grained C ontrastive L earning (UTE-FCL) to achieve a simple but efficient model. Specifically, in each recipe, UTE-FCL first concatenates each of the Ingredient and Instruction texts composed of multiple sentences as a single text. Then, it connects these two concatenated texts with the original single-phrase Title to obtain the concatenated recipe. Finally, it encodes these three concatenated texts and the original Title by a Transformer-based Unified Text Encoder (UTE). This proposed structure greatly reduces the memory usage and improves the feature encoding efficiency. Further, we propose fine-grained contrastive learning objectives to capture the correspondence between recipe components and the image at Title, Ingredient, and Instruction levels by measuring the mutual information. Extensive experiments demonstrate the effectiveness of UTE-FCL compared to existing methods. Haruya Kyutoku, Keisuke Doman, Takahiro Komamizu, Ichiro Ide, Jiangbo Qian |
Knowl. Based Syst. | 5 |
| 2024 | Computational measurement of perceived pointiness from pronunciationabstractAbstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications. Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 4 |
| 2024 | Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 4 |
| 2023 | RecipeMeta: Metapath-enhanced Recipe Recommendation on Heterogeneous Recipe NetworkabstractRecipe is a set of instructions that describes how to make food. It can help people from the preparation of ingredients, food cooking process, etc. to prepare the food, and increasingly in demand on the Web. To help users find the vast amount of recipes on the Web, we address the task of recipe recommendation. Due to multiple data types and relationships in a recipe, we can treat it as a heterogeneous network to describe its information more accurately. To effectively utilize the heterogeneous network, metapath was proposed to describe the higher-level semantic information between two entities by defining a compound path from peer entities. Therefore, we propose a metapath-enhanced recipe recommendation framework, RecipeMeta, that combines GNN (Graph Neural Network)-based representation learning and specific metapath-based information in a recipe to predict User-Recipe pairs for recommendation. Through extensive experiments, we demonstrate that the proposed model, RecipeMeta, outperforms state-of-the-art methods for recipe recommendation. Jialiang Shi, Takahiro Komamizu, Keisuke Doman, Haruya Kyutoku, Ichiro Ide |
MMAsia | 5 |
| 2023 | Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMM (1) | 5 |
| 2023 | Discovering Phonesthemic Clusters in Readings of Kanji Characters toward Exploring Phonestheme in Japanese
Akira Yoshida, Chihaya Matsuhira, Hirotaka Kato, Takatsugu Hirayama, Takahiro Komamizu, Ichiro Ide |
PACLIC | 6 |
| 2022 | Detection of Birds in a 3D Environment Referring to Audio-Visual InformationabstractWe propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed. Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi |
AVSS | 2 |
| 2022 | A Novel Approach for Pill-Prescription Matching with GNN Assistance and Contrastive Learning
Trung Thanh Nguyen 0006, Hoang Dang Nguyen, Thanh-Hung Nguyen, Hieu H. Pham 0001, Ichiro Ide, Phi-Le Nguyen |
PRICAI (1) | 5 |
| 2021 | MMArt-ACM'21: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2021abstractThe International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and social media. The topics of the accepted papers cover an analytic topic on comic contents understanding to generative topics on image synthesis and conversion. The actual MMArt-ACM'21 Proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3460426. Min-Chun Hu 0001, Ichiro Ide, Kensuke Tobitani |
ICMR | 2 |
| 2021 | Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 3 |
| 2021 | Soft-Boundary Label Relaxation with class placement constraints for semantic segmentation of the railway environment
Yuki Furitsu, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Hiroki Mukojima, Nozomi Nagamine |
Pattern Recognit. Lett. | 4 |
| 2020 | LFIR2Pose: Pose Estimation from an Extremely Low-resolution FIR image SequenceabstractIn this paper, we propose a method for human pose estimation from a Low-resolution Far-InfraRed (LFIR) image sequence captured by a 16 × 16 FIR sensor array. Human body estimation from such a single LFIR image is a hard task. For training the estimation model, annotation of the human pose to the images is also a difficult task for human. Thus, we propose the LFIR2Pose model which accepts a sequence of LFIR images and outputs the human pose of the last frame, and also propose an automatic annotation system for the model training. Additionally, considering that the scale of human body motion is largely different among body parts, we also propose a loss function focusing on the difference. Through an experiment, we evaluated the human pose estimation accuracy with an original data set, and confirmed that human pose can be estimated accurately from an LFIR image sequence. Saki Iwata, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa |
ICPR | 4 |
| 2020 | Ω-GAN: Object Manifold Embedding GAN for Image Generation by Disentangling Parameters into Pose and Shape ManifoldsabstractIn this paper, we propose Object Manifold Embedding GAN (Ω-GAN) to generate images of variously shaped and arbitrarily posed objects from a noise variable sampled from a distribution defined over the pose and the shape manifolds in a vector space. We introduce Parametric Manifold Sampling to sample noise variables from a distribution over the pose manifold to conditionally generate object images in arbitrary poses by tuning the pose parameter. We also introduce Object Identity Loss for clearly disentangling the pose and shape parameters, which allows us to maintain the shape of the object instance when only the pose parameter is changed. Through evaluation, we confirmed that the proposed Ω-GAN could generate variously shaped object images in arbitrary poses by changing the pose and shape parameters independently. We also introduce an application of the proposed method for object pose estimation, through which we confirmed that the object poses in the generated images are accurate. Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ICPR | 3 |
| 2020 | Median-Shape Representation Learning for Category-Level Object Pose Estimation in Cluttered EnvironmentsabstractIn this paper, we propose an occlusion-robust pose estimation method of an unknown object instance in an object category from a depth image. In a cluttered environment, objects are often occluded mutually. For estimating the pose of an object in such a situation, a method that de-occludes the unobservable area of the object would be effective. However, there are two difficulties; occlusion causes the offset between the center of the actual object and its observable area, and different instances in a category may have different shapes. To cope with these difficulties, we propose a two-stage Encoder-Decoder model to extract features with objects whose centers are aligned to the image center. In the model, we also propose the Median-shape Reconstructor as the second stage to absorb shape variations in a category. By evaluating the method with both a large-scale virtual dataset and a real dataset, we confirmed the proposed method achieves good performance on pose estimation of an occluded object from a depth image. Hiroki Tatemichi, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Ayako Amma, Hiroshi Murase |
ICPR | 4 |
| 2020 | MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020abstractThe International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173 Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki |
ICMR | 2 |
| 2020 | CEA'20: The 12th Workshop on Multimedia for Cooking and Eating ActivitiesabstractThe 12th Workshop on Multimedia for Cooking and Eating Activities presents This overview introduces the aim of the CEA'20 workshop and the list of papers presented in the workshop. Ichiro Ide, Yoko Yamakata, Atsushi Hashimoto 0001 |
ICMR | 1 |
| 2020 | Imageability Estimation using Visual and Language FeaturesabstractImageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability. Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
ICMR | 3 |
| 2020 | Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 2 |
| 2020 | More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 3 |
| 2020 | Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 2 |
| 2019 | Exemplar-Based Pseudo-Viewpoint Rotation for White-Cane User Recognition from a 2D Human Pose SequenceabstractIn recent years, various facilities are equipped to support visually impaired people, but accidents caused by visual disabilities still occur. In this paper, to support the visually-impaired people in a public space, we aim to classify whether a pedestrian image sequence obtained by a surveillance camera is a white-cane user or not from the temporal transition of a human pose represented as 2D coordinates. However, since the appearance of the 2D pose varies largely depending on the viewpoint of the pose, it is difficult to classify them. So, in this paper, we propose a method to rotate the viewpoint of a pose from various pseudo-viewpoints based on a pair of 2D poses simultaneously observed and classify the sequence by multiple classifiers corresponding to each viewpoint. Viewpoint rotation makes it possible to obtain pseudo-poses seen from various pseudo-viewpoints, extract richer pose features, and recognize white-cane users more accurately. Through an experiment, we confirmed that the proposed method improves the recognition rate by 12% compared to the method not employing viewpoint rotation. Naoki Nishida 0003, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Jun Piao |
AVSS | 4 |
| 2019 | Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 2 |
| 2018 | Gaze-Inspired Learning for Estimating the Attractiveness of a Food PhotoabstractThe number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them. Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ISM | 5 |
| 2018 | Voting-based Hand-Waving Gesture Spotting from a Low-Resolution Far-Infrared Image SequenceabstractWe propose a temporal spotting method of a hand gesture from a low-resolution far-infrared image sequence captured by a far-infrared sensor array. The sensor array captures the spatial distribution of far-infrared intensity as a thermal image by detecting far-infrared waves emitted from heat sources. It is difficult to spot a hand gesture from a sequence of thermal images captured by the sensor due to its low-resolution, heavy noise, and varying duration of the gesture. Therefore, we introduce a voting-based approach to spot the gesture with template matching-based gesture recognition. We confirm the effectiveness of the proposed temporal spotting method in several settings. Yasutomo Kawanishi, Chisato Toriyama, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa, Masato Kawade |
VCIP | 5 |
| 2017 | Action recognition from extremely low-resolution thermal image sequenceabstractThis paper proposes a Deep Learning-based action recognition method from an extremely low-resolution thermal image sequence. The method recognizes daily actions by humans (e.g. walking, sitting down, standing up, etc.) and abnormal actions (e.g. falling down) without privacy concerns. While privacy concerns can be ignored, it is difficult to compute feature points and to obtain a clear edge of the human body from an extremely low-resolution thermal image. To address these problems, this paper proposes a Deep Learning-based action recognition method that combines convolution layers and an LSTM layer for learning spatio-temporal representation, whose inputs are the thermal images and their frame differences cropped by the gravity center of human regions. The effectiveness of the proposed method was confirmed through experiments. Takayuki Kawashima, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Daisuke Deguchi, Tomoyoshi Aizawa, Masato Kawade |
AVSS | 3 |
| 2017 | Automatic Selection of Web Contents Towards Automatic Authoring of a Video BiographyabstractIn this paper, we propose a method for image selection using Web image search for automatic video biography authoring. In the proposed method, images are selected from the image search results considering their visual contents for inclusion in the video biography. Through evaluation, we confirmed the effectiveness of the proposed image selection method compared to a baseline method which simply selects the top 1 search result. Ichiro Ide, Yasutomo Kawanishi, Kyoka Kunishiro, Frank Nack, Daisuke Deguchi, Hiroshi Murase |
ISM | 1 |
| 2017 | Summarization of News Videos Considering the Consistency of Auditory and Visual ContentsabstractSince news videos are valuable sources of multimedia information on real-world events, there is a demand for viewing them efficiently. However, there is a problem that summarization methods based on auditory contents do not take into account the visual contents. In the case of news videos, due to its presentation style where audio contents and visual contents do not necessarily come from the same source, this could severely decrease the amount of informative visual contents included in the generated summarized video. Thus, we propose a method for summarizing a sequence of news videos considering the consistency of both auditory and visual contents. The proposed method first selects key-sentences from the auditory contents (Closed Caption) of each news story in the sequence, and then selects a shot within the news story whose "Visual Concepts" detected from the visual contents are the most consistent with the key-phrase. Finally, the audio segment corresponding to each key-phrase is overlapped onto the selected shot, and then concatenated to generate a summarized video. The effectiveness of the proposed method was confirmed on several news topics through a subjective experiment. Ichiro Ide, Ryunosuke Tanishige, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase |
ISM | 1 |
| 2017 | Monocular localization within sparse voxel mapsabstractWe introduce a method that uses a single camera to localize a vehicle within a pre-constructed map consisting of a voxel occupancy grid and road-line marker positions. Sophisticated mapping hardware is capable of creating high-accuracy 3D maps of road environments, but localizing a vehicle within such maps is one of the challenges at the forefront of automated driving. A solution which is robust to dynamic environments, while using only inexpensive sensors, is a difficult problem. In addition, maps that enable precise localization consume a lot of data which is impractical for the expansive environments encountered in real-world road networks. We show how using the area of edge regions shared between rendered views of a compact voxel map and in-vehicle camera images can be coupled with non-linear optimization methods to determine the camera position and pose. David Wong 0002, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2017 | Proposal of a spectral random dots marker using local feature for posture estimationabstractWe propose a novel marker for robot's grasping task which has the following three aspects: (i) it is easy-to-find in a cluttered background, (ii) it is calculable for its posture (iii) its size is compact. The proposed marker is composed of a random dots pattern, and uses keypoint detection and a scale estimation by Spectral SIFT for dots detection and data decoding. The data is encoded by the scale size of dots, and the same dots in the marker work for both marker detection and data decoding. As a result, the proposed marker size can be compact. We confirmed the effectiveness of the proposed marker through experiments. Norimasa Kobori, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
VR | 3 |
| 2017 | Regression of feature scale tracklets for decimeter visual localization
David Wong 0002, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase |
Image Vis. Comput. | 4 |
| 2016 | A classification method of cooking operations based on eye movement patternsabstractWe are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables. Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ETRA | 5 |
| 2016 | Moving camera background-subtraction for obstacle detection on railway tracksabstractWe propose a method for detecting obstacles by comparing input and reference train frontal view camera images. In the field of obstacle detection, most methods employ a machine learning approach, so they can only detect pre-trained classes, such as pedestrian, bicycle, etc. This means that obstacles of unknown classes cannot be detected. To overcome this problem, we propose a background subtraction method that can be applied to moving cameras. First, the proposed method computes frame-by-frame correspondences between the current and the reference (database) image sequences. Then, obstacles are detected by applying image subtraction to corresponding frames. To confirm the effectiveness of the proposed method, we conducted an experiment using several image sequences captured on an experimental track. Its results showed that the proposed method could detect various obstacles accurately and effectively. Hiroki Mukojima, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Masato Ukai, Nozomi Nagamine, Ryuta Nakasone |
ICIP | 4 |
| 2016 | Misclassification tolerable learning for robust pedestrian orientation classificationabstractIn this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset. Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
ICPR | 3 |
| 2016 | Parts Selective DPM for detection of pedestrians possessing an umbrellaabstractIn recent years, pedestrian detection from an in-vehicle camera has been attracting attention. However, in the case of a raining situation, the detection accuracy decreases because the head of a pedestrian tends to be occluded by an umbrella. In oder to handle such cases, in this paper, as a variation of the Deformable Part Model (DPM) which is widely used in the field of object recognition, we propose “Parts Selective DPM (PS-DPM)” which selectively chooses the original part filters and additional part filters trained independently. In the detection of pedestrians possessing an umbrella, the selection of head and umbrella parts will make pedestrian detection more robust to the occlusion. We conducted experiments to evaluate the performance of the proposed method. As a result, pedestrian detection with the proposed PS-DPM achieved high detection accuracy in rainy weather, compared with the detection by the conventional DPM. Moreover, we confirmed that it did not decrease the pedestrian detection accuracy in fine weather. Yuto Shimbo, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2015 | Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF cameraabstractThis paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification. Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
Intelligent Vehicles Symposium | 4 |
| 2014 | Spatial People Density Estimation from Multiple Viewpoints by Memory Based RegressionabstractCrowd analysis using cameras has attracted much attention for public safety and marketing. Among techniques of the crowd analysis, we focus on spatial people density estimation which estimates the number of people for each small area in a floor region. However, spatial people density cannot be estimated accurately for an area far from the camera because of the occlusion by people in a closer area. Therefore, we propose a method using a memory based regression method with images captured from cameras from multiple viewpoints. This method is realized by looking up a table that consists of correspondences between people density maps and crowd appearances. Since the crowd appearances include situations where various occlusions occur, an estimation robust to occlusion should be realized. In an experiment, we examined the effectiveness of the proposed method. Yoshimune Tabuchi, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takayuki Kurozumi, Kunio Kashino |
ICPR | 4 |
| 2014 | Scene Duplicate Detection from News Videos Using Image-Audio Matching Focusing on Human FacesabstractAs one tool for structuring a massive volume of archived news videos based on their semantic contents, this paper proposes a method to detect scene duplicates from news videos. A scene duplicate is a pair of video segments taken at the same event from different viewpoints. Referring to the audio channel is effective to detect scene duplicates regardless of viewpoints, but it cannot be relied on when external audio sources (e.g. Narrations, sound effects) overlap the original one. In contrast, the image channel can be useful in most cases, although significant difference in viewpoints affect the detection. The proposed method integrates the information from these two channels in order to improve the accuracy of scene duplicate detection from news videos. The performance of the proposed method was evaluated through an experiment with actual broadcast news videos. As a result, we obtained the higher detection accuracies in both recall and precision. Therefore, we confirmed the effectiveness of the proposed method. Haruka Kumagai, Keisuke Doman, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ISM | 3 |
| 2014 | Estimation of traffic sign visibility considering local and global features in a driving environmentabstractThis paper proposes a camera-based visibility estimation method for a traffic sign. The visibility here indicates how a visual target is easy to be detected and recognized by a human driver (not a machine). This research aims at realizing a nuisance-free driver assistance system which sorts out information depending on the visibility of a visual target, in order to prevent driver distraction. Our previous study on estimating the visibility of a traffic sign considered only the effect of the local region around a target, assuming the situation that a driver's gaze is around it. The proposed method integrates both the local features and global features in a driving environment without such an assumption. The global features evaluate the positional relationships between traffic signs and the appearance around the fixation point of a driver's gaze, which considers the effect of the driver's entire field of view. Experimental results showed the effectiveness of incorporating the global features for estimating the visibility of a traffic sign. Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Utsushi Sakai |
Intelligent Vehicles Symposium | 5 |
| 2014 | Single camera vehicle localization using SURF scale and dynamic time warpingabstractVehicle ego-localization is an essential process for many driver assistance and autonomous driving systems. The traditional solution of GPS localization is often unreliable in urban environments where tall buildings can cause shadowing of the satellite signal and multipath propagation. Typical visual feature based localization methods rely on calculation of the fundamental matrix which can be unstable when the baseline is small. In this paper we propose a novel method which uses the scale of matched SURF image features and Dynamic Time Warping to perform stable localization. By comparing SURF feature scales between input images and a pre-constructed database, stable localization is achieved without the need to calculate the fundamental matrix. In addition, 3D information is added to the database feature points in order to perform lateral localization, and therefore lane recognition. From experimental data captured from real traffic environments, we show how the proposed system can provide high localization accuracy relative to an image database, and can also perform lateral localization to recognize the vehicle's current lane. David Wong 0002, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 3 |
| 2014 | Estimation of the Representative Story Transition in a Chronological Semantic Structure of News TopicsabstractIt is important to track the flow of topics to thoroughly understand the contents. Accordingly, a method that structures the chronological semantic relations between news stories, namely a "topic thread structure" has been proposed. It allows the comprehensive understanding of a topic by chronologically tracking stories one by one from the initial story. However, this task imposes a user to watch many stories when it contains various sub-topics. Thus, we propose a method that estimates the representative story transition in a topic thread structure. In the proposed method, features obtained from a story and those from the topic thread structure are used for the estimation. We confirmed the effectiveness of the proposed method by comparing the results obtained from the proposed method to the ground truth obtained from votes in a subjective experiment. Kosuke Kato, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ICMR | 2 |
| 2014 | Event Detection based on Twitter Enthusiasm Degree for Generating a Sports Highlight VideoabstractThis paper presents a Twitter-based event detection method based on "Twitter Enthusiasm Degrees (TED)" toward generating a highlight video of a sports game. Existing methods not only depend on both languages and sports types but also often falsely detect non-target events. In contrast, the proposed method detects sports events using TEDs calculated from several kinds of string features independent of languages and sports. We applied the proposed method to actual sports games, and compared the detected events with the events present in broadcasted highlight videos, and confirmed the effectiveness and the language and sports type independencies of the proposed method. Keisuke Doman, Taishi Tomita, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ACM Multimedia | 3 |
| 2013 | Pedestrian detection by scene dependent classifiers with generative learningabstractRecently, pedestrian detection from in-vehicle camera images is becoming an crucial technology for Intelligent Transportation Systems (ITS). However, it is difficult to detect pedestrians accurately in various scenes by obtaining training samples. To tackle this problem, we propose a method to construct scene dependent classifiers to improve the accuracy of pedestrian detection. The proposed method selects an appropriate classifier based on the scene information that is a category of appearance associated with location information. To construct scene dependent classifiers, the proposed method introduces generative learning for synthesizing scene dependent training samples. Experimental results showed that the detection accuracy of the proposed method outperformed the comparative method, and we confirmed that scene dependent classifiers improved the accuracy of pedestrian detection. Hidefumi Yoshida, Daichi Suzuo, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takashi Machida, Yoshiko Kojima |
Intelligent Vehicles Symposium | 4 |
| 2013 | Recompilation of Broadcast Videos Based on Real-World Scenarios
Ichiro Ide |
MMM (2) | 1 |
| 2013 | Detection of Biased Broadcast Sports Video Highlights by Attribute-Based Tweets Analysis
Takashi Kobayashi 0001, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
MMM (2) | 4 |
| 2012 | Robust Face Super-Resolution Using Free-Form Deformations for Low-Quality Surveillance VideoabstractRecently, the demand for face recognition to identify persons from surveillance video cameras has rapidly increased. Since surveillance cameras are usually placed at positions far from a person's face, the quality of face images captured by the cameras tends to be low. This degrades the recognition accuracy. Therefore, aiming to improve the accuracy of the low-resolution-face recognition, we propose a video-based super-resolution method. The proposed method can generate a high-resolution face image from low-resolution video frames including non-rigid deformations caused by changes of face poses and expressions without using any positional information of facial feature points. Most existing techniques use the facial feature points for image alignment between the video frames. However, it is difficult to obtain the accurate positions of the feature points from low-resolution face images. To achieve the alignment, the proposed method uses a free-form deformation method that flexibly aligns each local region between the images. This enables super-resolution of face images from low-resolution videos. Experimental results demonstrated that the proposed method improved the performance of super-resolution for actual videos in terms of both image quality and face recognition accuracy. Tomonari Yoshida, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ICME | 4 |
| 2012 | Estimation of the human performance for pedestrian detectability based on visual search and motion features
Masashi Wakayama, Daisuke Deguchi, Keisuke Doman, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu |
ICPR | 4 |
| 2012 | Smart VideoCooKing: a multimedia cooking recipe browsing application on portable devicesabstractThis demo presents "Smart VideoCooKing" which is a multimedia cooking recipe browsing application on portable Android devices. A multimedia cooking recipe is a cooking recipe where each cooking operation is associated with a corresponding video clip describing it, aimed to facilitate the understanding of cooking operations. In combination with third-party applications, "Smart VideoCooKing" provides useful functions such as playing cooking video clips describing cooking operations quickly, searching information of ingredients easily, and reading aloud cooking directions. Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
ACM Multimedia | 4 |
| 2012 | Overview of the ACM multimedia 2012 workshop on multimedia for cooking and eating activities (CEA'12)abstractThis overview introduces the aim of the CEA'12 workshop and the list of papers presented in the workshop. Mutsuo Sano, Ichiro Ide, Yoko Yamakata |
ACM Multimedia | 2 |
| 2011 | Low Resolution QR-Code Recognition by Applying Super-Resolution Using the Property of QR-CodesabstractThis paper proposes a method for low resolution QR-code recognition. A QR-code is a two-dimensional binary symbol that can embed various information such as characters and numbers. To recognize a QR-code correctly and stably, the resolution of an input image should be high. In practice, however, recognition of a QR-code is usually difficult due to low resolution when it is captured from a distance. In this paper, we propose a method to improve the performance of low resolution QR-code recognition by using the super-resolution technique that generates a high resolution image from multiple low-resolution images. Although a QR-code is a binary pattern, it is observed as a grayscale image due to the degradation through the capturing process. Especially the pixels around the borders between white and black regions become ambiguous. To overcome this problem, the proposed method introduces a binary pattern constraint to generate super-resolved images appropriate for recognition. Experimental results showed that a recognition rate of 98% can be achieved by the proposed method, which is a 15.7% improvement in comparison with a method using a conventional super-resolution method. Yuji Kato, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
ICDAR | 4 |
| 2011 | Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News VideosabstractWe propose a method to detect the inconsistency between a subject and the speaker for extracting speech scenes from news videos. Speech scenes in news videos contain a wealth of multimedia information, and are valuable as archived material. In order to extract speech scenes from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such approach, since news videos contain non-speech scenes where the speaker is not the subject, such as narrated scenes. To solve this problem, we propose a method to discriminate between speech scenes and narrated scenes based on the co-occurrence between a subject's lip motion and the speaker's voice. The proposed method uses lip shape and degree of lip opening as visual features representing a subject's lip motion, and uses voice volume and phoneme as audio feature representing a speaker's voice. Then, the proposed method discriminates between speech scenes and narrated scenes based on the correlations of these features. We report the results of experiments on videos captured in a laboratory condition and also on actual broadcast news videos. Their results showed the effectiveness of our method and the feasibility of our research goal. Shogo Kumagai, Keisuke Doman, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ISM | 5 |
| 2011 | Intelligent traffic sign detector: Adaptive learning based on online gathering of training samplesabstractThis paper proposes an intelligent traffic sign detector using adaptive learning based on online gathering of training samples from in-vehicle camera image sequences. To detect traffic signs accurately from in-vehicle camera images, various training samples of traffic signs are needed. In addition, to reduce false alarms, various background images should also be prepared before constructing the detector. However, since their appearances vary widely, it is difficult to obtain them exhaustively by manual intervention. Therefore, the proposed method simultaneously obtains both traffic sign images and background images from in-vehicle camera images. Especially, to reduce false alarms, the proposed method gathers background images that were easily mis-detected by a previously constructed traffic sign detector, and re-trains the detector by using them as negative samples. By using retrospectively tracked traffic sign images and background images as positive and negative training samples, respectively, the proposed method constructs a highly accurate traffic sign detector automatically. Experimental results showed the effectiveness of the proposed method. Daisuke Deguchi, Daisuke Shirasuna, Keisuke Doman, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2011 | Estimation of traffic sign visibility considering temporal environmental changes for smart driver assistanceabstractWe propose a visibility estimation method for traffic signs considering temporal environmental changes, as a part of work for the realization of nuisance-free driver assistance systems. Recently, the number of driver assistance systems in a vehicle is increasing. Accordingly, it is becoming important to sort out appropriate information provided from them, because providing too much information may cause driver distraction. To solve such a problem, we focus on a visibility estimation method for controlling the information according to the visibility of a traffic sign. The proposed method sequentially captures a traffic sign by an in-vehicle camera, and estimates its accumulative visibility by integrating a series of instantaneous visibility. By this way, even if the environmental conditions may change temporally and complicatedly, we can still accurately estimate the visibility that the driver perceives in an actual traffic scene. We also investigate the performance of the proposed method and show its effectiveness. Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu |
Intelligent Vehicles Symposium | 5 |
| 2011 | Road image update using in-vehicle camera images and aerial imageabstractRoad image is becoming important for several applications such as car navigation systems, traffic environment research, city modeling. Usually, a road image can be obtained from an aerial image but the resolution of the aerial image is often low, or it contains occlusions by obstacles. Therefore, the update of road image is required. In this paper, we propose a road image mosaicing method using in-vehicle camera images and an aerial image. We first perform image registration of road regions between these images, and then, we generate a large road image by performing image mosaicing of road regions in invehicle camera images. In an experiment, we achieved resolution improvement and occlusions removal, and also succeeded in update of a large road image. Masafumi Noda, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Yoshiko Kojima, Takashi Naito |
Intelligent Vehicles Symposium | 4 |
| 2011 | 3-D line segment reconstruction using an in-vehicle camera for free space detectionabstractFree space detection is very important for vehicle navigation and safe driving. 3-D line segment reconstruction of a street is important for the free space detection because a street-view includes many line segments. For the free space detection, we propose a method for reconstructing 3-D line segments in a streetscape using a monocular in-vehicle camera. The 3-D reconstruction of the line segments is achieved by using each three images from an image sequence. Once accurate camera poses of these images are obtained, one of the remaining crucial problems is to match the line segments between the images correctly. A strategy for finding correspondence of the line segments is as follows: First, the correspondences of line segment candidates are searched by using a two-view constraint. However, the two-view constraint has difficulty on determining an unique correspondence geometrically. Therefore, the candidates of the line segment correspondences are reduced using a three-view constraint. In order to improve the accuracy, the proposed method exploits a color feature of the line segment and a preliminary knowledge of the vehicle motion. Finally, the line segments are reconstructed using the correspondences. From an experimental result, we confirmed the effectiveness of the proposed method. Application to the free space detection demonstrated the usefulness of the reconstructed line segments. Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2011 | Why did the prime minister resign?: generation of event explanations from large news repositoriesabstractOne of the common parts of news is to provide the background for a current event, such as the resignation of a Prime Minister. This paper addresses a framework that facilitates semi-automated authoring of explanatory audio-visual news topics in a retrospective style for the domain of politics based on already edited new stories available in the repository of the news corporation. The aim is to facilitate a journalist with an audio-visual body based on which he/she can finalize the explanatory piece. The proposed framework enhances current state of the art video summarization by allowing the combination of different news stories into one coherent explanation about a topic of the current news. The framework introduces techniques that exploit demoscopic data in form of polls for the development of the general story outline; the automatic retrieval of relevant material by using a combination of event templates and automatic news summarization over topic threads; and the generation of the final video by applying a set of trimming rules. Example generations are presented and discussed and an outline of future work is presented. Frank Nack, Ichiro Ide |
ACM Multimedia | 2 |
| 2011 | Scene segmentation of wedding party videos by scenario-based matching with example videosabstractWe propose a method for scene segmentation of a wedding party video. Recently, it has become popular to take videos of a wedding ceremony and its party. Especially, because of its length, each scene of a wedding party video needs to be indexed with each event for efficient browsing. The proposed method segments a wedding party video into scenes of events by scenario- based matching with example videos that are synthesized by combining scenes from other wedding party videos according to a scenario. Kazuki Sawai, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ACM Multimedia | 4 |
| 2011 | Video CooKing: Towards the Synthesis of Multimedia Cooking Recipes
Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
MMM (2) | 4 |
| 2010 | Classification of Near-Duplicate Video Segments Based on Their Appearance PatternsabstractWe propose a method that analyzes the structure of a large volume of general broadcast video data by the appearance patterns of near-duplicate video segments. We define six classification rules based on the appearance patterns of near-duplicate video segments according to their roles, and evaluated them over more than 1,000 hours of actual broadcast video data. Ichiro Ide, Yuji Shamoto, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase |
ICPR | 1 |
| 2010 | Region-Based Image Transform for Transition Between Object AppearancesabstractWe propose a method of region-based image transform to achieve accurate transition between object appearances. A view-transition model (VTM) is one of the statistical methods that learn appearance transition from a sample image dataset of a large number of objects with various appearances. However, the VTM method has a practical problem that the appearance transition cannot be performed accurately if a sufficient number of learning samples is not available in the dataset. To cope with the problem, the proposed method first determines the regions of input and output images whose pixel values mutually affect each other during appearance transition, then transforms iteratively between partial images in the regions. We conducted experiments using actual image datasets. The results show that the proposed method could accurately transform appearances compared with the VTM method. Tomokazu Takahashi, Yuki Kono, Ichiro Ide, Hiroshi Murase |
ICPR | 3 |
| 2010 | Removal of Moving Objects from a Street-View Image by Fusing Multiple Image SequencesabstractWe propose a method to remove moving objects from an in-vehicle camera image sequence by fusing multiple image sequences. Driver assistance systems and services such as Google Street View require images containing no moving object. The proposed scheme consists of three parts: (i) collection of many image sequences along the same route by using vehicles equipped with an omni-directional camera, (ii) temporal and spatial registration of image sequences, and (iii) mosaicing partial images containing no moving object. Experimental results show that 97.3% of the moving object area could be removed by the proposed method. Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
ICPR | 4 |
| 2010 | Multimedia Supplementation to a Cooking Recipe Text for Facilitating Its Understanding to Inexperienced UsersabstractAssisting culinary activities for inexperienced users has been considered as an important task in most existing works in the field. On the other hand, recipe texts are becoming available on the Internet in increasing numbers. However, they tend to be written simply by mostly non-professional people, and thus are sometimes difficult for an inexperienced person to follow the steps and manage to cook as they are supposed to. In this paper, we propose a method that detects difficult descriptions for an inexperienced user in an existing text recipe, and supplements them with multimedia contents including text information extracted from a large number of recipes, and also images and video clips on certain kinds of cooking operations, to facilitate the understanding of the recipe. Experimental results showed promising ability of the proposed method to assist inexperienced users understand the descriptions in a recipe. Ichiro Ide, Yuka Shidochi, Yuichi Nakamura 0001, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase |
ISM | 1 |
| 2010 | Estimation of traffic sign visibility toward smart driver assistanceabstractWe propose a visibility estimation method for traffic signs as part of work for realization of nuisance-free driving safety support systems. Recently, the number of driving safety support systems in a car has been increasing. As a result, it is becoming important to select appropriate information from them for safe and comfortable driving because too much information may cause driver distraction and may increase the risk of a traffic accident. One of the approaches to avoid such a problem is to alert the driver only with information which could easily be missed. Therefore, to realize such a system, we focus on estimating the visibility of traffic signs. The proposed method is a model-based method that estimates the visibility of traffic signs focusing on the difference of image features between a traffic sign and its surrounding region. In this paper, we investigate the performance of the proposed method and show its effectiveness. Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu |
Intelligent Vehicles Symposium | 5 |
| 2010 | PageRank with Text Similarity and Video Near-Duplicate Constraints for News Story Re-ranking
Xiaomeng Wu, Ichiro Ide, Shin'ichi Satoh 0001 |
MMM | 2 |
| 2010 | A Hilbert warping method for handwriting gesture recognition
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
Pattern Recognit. | 3 |
| 2009 | Low-Resolution Character Recognition by Video-Based Super-ResolutionabstractIn this paper, we propose a method for recognizing low-resolution characters using a super-resolution technique. Although portable digital cameras can be used for camera based character recognition, the captured images contain several types of noises which make the recognition task difficult. We introduce a phase of super-resolution before the recognition to enhance the resolution of images obtained from a video. The proposed method uses the subspace method for the recognition of characters which are integrated from multiple low-resolution characters by the super-resolution technique. Experimental results show that the proposed method improves the recognition accuracy; we confirmed that the recognition rate for the input size of 7 times 7 pixels was 90.35%, and for the input size of 9 times 9 pixels was 99.97%. Ataru Ohkura, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
ICDAR | 4 |
| 2009 | Adaptive division of feature space for rapid detection of near-duplicate video segmentsabstractNear-duplicate video detection is becoming a core-technology for analyzing the structure of a large-scale video archive. It, however, is naturally an O(n2) problem, where n is a value proportional to the total length of an input video stream. We have previously challenged this time-consuming task by reducing the cost required for each of the O(n2) comparisons. This paper, on the other hand, proposes a method that reduces the number of comparisons by adaptively dividing the feature space according to the distribution of feature points. Ichiro Ide, Shugo Suzuki, Tomokazu Takahashi, Hiroshi Murase |
ICME | 1 |
| 2009 | Labeling News Topic Threads with Wikipedia EntriesabstractWikipedia is a famous online encyclopedia. However most Wikipedia entries are mainly explained by text, so it will be very informative to enhance the contents with multimedia information such as videos. Thus we are working on a method to extend information of Wikipedia entries by means of broadcast videos which explain the entries. In this work, we focus especially on news videos and Wikipedia entries about news events. In order to extend information of Wikipedia entries, it is necessary to link news videos and Wikipedia entries. So the main issue will be on a method that labels news videos with Wikipedia entries automatically. In this way, explanations could be more detailed with news videos can be exhibited, and the context of the news events should become easier to understand. Through experiments, news videos were accurately labeled with Wikipedia entries with a precision of 86% and a recall of 79%. Tomoki Okuoka, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ISM | 4 |
| 2009 | Overview of the ACM multimedia 2009 workshop on multimedia for cooking and eating activities (CEA'09)abstractThis overview introduces the aim of the CEA'09 workshop and the list of papers presented in the workshop. Mutsuo Sano, Ichiro Ide, Kenzaburo Miyawaki |
ACM Multimedia | 2 |
| 2009 | A Multimodal Constellation Model for Object Category Recognition
Yasunori Kamiya, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
MMM | 3 |
| 2008 | A Hilbert Warping Algorithm for Recognizing Characters from Moving CameraabstractWe present a method for recognizing characters from image sequences captured by moving camera. In the proposed method, the sequence of the captured images is compared with those of reference character patterns using the concept of analytic signal. Since the captured image sequence can be nonlinearly warped along the time axis due to the movement of a hand-held camera, phase synchronization of two analytic signals is used for the alignment of two image sequences. Hilbert transform is used to convert all the image sequences into analytic signals whose phases are supposed to be increasing. Experimental results showed the usefulness of the proposed phase-based alignment algorithm. Hiroyuki Ishida, Ichiro Ide, Hiroshi Murase, Tomokazu Takahashi |
Document Analysis Systems | 2 |
| 2008 | A Hilbert warping method for camera-based finger-writing recognitionabstractWe propose a time-warping algorithm for recognizing finger actions by a camera. In the proposed method, an input image sequence is aligned to the reference sequences by phase-synchronization of the analytic signals, and then classified by comparing the cumulative distances. A major benefit of this method is that over-fitting to sequences of incorrect categories is restricted. The proposed method exhibited high recognition accuracy in finger-writing character recognition. Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
ICPR | 3 |
| 2008 | Eigenspace interpolation for appearance-based object recognitionabstractAn eigenspace interpolation method smoothly interpolates between two different eigenspaces using high dimensional rotation. However, up to now its effectiveness in object recognition and the validity of the interpolation algorithm have not been discussed sufficiently. We therefore propose an appearance-based object recognition method combining the eigenspace interpolation method and a subspace method. We conducted face recognition experiments using images captured from multiple camera positions with various illumination conditions. Experimental results demonstrate the effectiveness of the proposed method and the validity of the interpolation algorithm. Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase |
ICPR | 3 |
| 2008 | Cross-Lingual Retrieval of Identical News Events by Near-Duplicate Video Segment Detection
Akira Ogawa, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase |
MMM | 3 |
| 2008 | Recognition of camera-captured low-quality characters using motion blur information
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Yoshito Mekada, Hiroshi Murase |
Pattern Recognit. | 3 |
| 2007 | Interpolation Between Eigenspaces Using Rotation in Multiple Dimensions
Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase |
ACCV (2) | 3 |
| 2007 | Genre-Adaptive Near-Duplicate Video Segment DetectionabstractThis paper proposes a fast and accurate method to detect all near-duplicate segments in a video stream. To reduce the computation time while ensuring the detection accuracy equivalent to that by brute-force frame-by-frame comparison, a two-step detection method is proposed; a fast but rough detection applied in a compressed feature vector space spanned by the result of a PC A, followed by confirmation of candidates in the original high dimension space. The results show that the proposed method accelerates the detection by more than 1,000 times while maintaining the detection accuracy. We also propose an entropy-based pixel selection scheme to generate feature vectors optimized for comparison of video segments within programs with mostly common pictures. The results show that the proposed scheme eliminates the false positives drastically, which should lead to even faster detection. Ichiro Ide, Kazuhiro Noda, Tomokazu Takahashi, Hiroshi Murase |
ICME | 1 |
| 2007 | mediaWalker: a video archive explorer based on time-series semantic structureabstractWe introduce a video browsing interface 'mediaWalker' that lets users explore a news video archive based on a time-series semantic structure; the 'topic thread' structure. The interface lets users efficiently track up and down the development of news in an archive with more than 1,000 hours of video. Ichiro Ide, Tomoyoshi Kinoshita, Tomokazu Takahashi, Shin'ichi Satoh 0001, Hiroshi Murase |
ACM Multimedia | 1 |
| 2006 | Spatiotemporal Density Feature Analysis to Detect Liver Cancer from Abdominal CT Angiography
Yoshito Mekada, Yuki Wakida, Yuichiro Hayashi, Ichiro Ide, Hiroshi Murase |
ACCV (2) | 4 |
| 2006 | Exploiting Topic Thread Structures in a News Video Archive for the Semi-Automatic Generation of Video SummariesabstractWe propose a method that semi-automatically composes video stories by connecting individual stories in a news video archive along a topic-based semantic structure, namely the topic thread. We introduce the methods to realize the composition, namely, story segmentation, topic threading and clustering. We then evaluate the proposed approach based on preliminary tests. Since the thread structure reflects the development of topics in the real-world, we believe that the composed news video story should be effective for the user to gain a deeper understanding of the current topic of interest Ichiro Ide, Hiroshi Mo, Norio Katayama, Shin'ichi Satoh 0001 |
ICME | 1 |
| 2005 | Automated Nomenclature of Bronchial Branches Extracted from CT Images and Its Application to Biopsy Path Planning in Virtual Bronchoscopy
Kensaku Mori, Sinya Ema, Takayuki Kitasaka, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori |
MICCAI (2) | 5 |
| 2005 | Cooking navi: assistant for daily cooking in kitchenabstractWe are developing a cooking navigation system, which helps even a novice user to cook several recipes in parallel without failure, while improving an advanced user's skill further. To realize this, the system optimizes the cooking procedure considering the following restrictions: (1) Duration of cooking, (2) Accuracy of cooking, and (3) Learning effect, by providing appropriate instructions to user's at the right timing, making full use of multimedia information. The users should be able to cook perfectly and comfortably just by following the text, video and audio provided by the system. According to the result of a preliminary experiment, all users from novice to experienced cooks could finish two dishes in parallel while enjoyeing the cooking very much. The result of a questionnaire shows the effectiveness of the multimedia navigation that we propose. Reiko Hamada, Jun Okabe, Ichiro Ide, Shin'ichi Satoh 0001, Shuichi Sakai, Hidehiko Tanaka |
ACM Multimedia | 3 |
| 2003 | Topic-based inter-video structuring of a large-scale news video corpusabstractWe propose a topic-based inter-video news video corpus structuring method and a visual interface to efficiently browse through the structured corpus. Such inter-video structuring was not deeply sought in previous works. The topic-based structure is analyzed by closed-caption text analysis; topic segmentation and tracking. The visual interface provides the ability to 1) search and select a topic by query terms and 2) track a topic thread interactively referring to the text analysis results. Although topic retrieval is somewhat similar to conventional video retrieval methods, the combination with topic tracking makes it remarkably easy to narrow down the results that match a user's interest and moreover reveal underlying content-based structures, where the structure itself contains rich information. Ichiro Ide, Hiroshi Mo, Norio Katayama, Shin'ichi Satoh 0001 |
ICME | 1 |
| 2002 | An object detection method for describing soccer games from videoabstractWe propose a novel object detection and tracking method in order to detect and track objects necessary to describe contents of a soccer game. On the contrary to intensity oriented conventional object detection methods, the proposed method refers to color rarity and local edge property, and integrally evaluates them by a fuzzy function to achieve better detection quality. These image features were chosen considering the characteristics of soccer video images, that most non-object regions are roughly single colored (green) and most objects tend to have locally strong edges. We also propose a simple object tracking method, that could track objects with occlusion with other objects using a color based template matching. The result of an evaluation experiment applied to actual soccer video showed very high detection rate in detecting player regions without occlusion, and promising ability for regions with occlusion. Okihisa Utsumi, Koichi Miura, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka |
ICME (1) | 3 |
| 1999 | Associating video with related documentsabstractArticle Free Access Share on Associating video with related documents Authors: Reiko Hamada Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Ichiro Ide Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Shuichi Sakai Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile , Hidehiko Tanaka Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan Graduate School of Electrical Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, JapanView Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999Pages 17–20https://doi.org/10.1145/319878.319883Published:01 October 1999Publication History 1citation199DownloadsMetricsTotal Citations1Total Downloads199Last 12 Months11Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Reiko Hamada, Ichiro Ide, Shuichi Sakai, Hidehiko Tanaka |
ACM Multimedia (2) | 2 |