VLDB 2026 Research / reviewers in the wild / expert
Takatsugu Hirayama
dblp:94/4172
· DBLP profile ↗
35ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0001-6290-9680ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying Image-Adjective Associations by Leveraging Large-Scale Pretrained Models
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide |
MMM (4) | 4 |
| 2025 | Pre-Instruction for Pedestrians Interacting Autonomous Vehicles With eHMI: Effects on Their Psychology and Walking Behavior
Hailong Liu 0001, Takatsugu Hirayama |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image GenerationabstractText-to-image diffusion models sometimes depict blended concepts in the generated images. One promising use case of this effect would be the nonword-to-image generation task which attempts to generate images intuitively imaginable from a non-existing word (nonword). To realize nonword-to-image generation, an existing study focused on associating nonwords with similar-sounding words. Since each nonword can have multiple similar-sounding words, generating images containing their blended concepts would increase intuitiveness, facilitating creative activities and promoting computational psycholinguistics. Nevertheless, no existing study has quantitatively evaluated this effect in either diffusion models or the nonword-to-image generation paradigm. Therefore, this paper first analyzes the conceptual blending in a pretrained diffusion model, Stable Diffusion. The analysis reveals that a high percentage of generated images depict blended concepts when inputting an embedding interpolating between the text embeddings of two text prompts referring to different concepts. Next, this paper explores the best text embedding space conversion method of an existing nonword-to-image generation framework to ensure both the occurrence of conceptual blending and image generation quality. We compare the conventional direct prediction approach with the proposed method that combines k-nearest neighbor search and linear regression. Evaluation reveals that the enhanced accuracy of the embedding space conversion by the proposed method improves the image generation quality, while the emergence of conceptual blending could be attributed mainly to the specific dimensions of the high-dimensional text embedding space. Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide |
ACM Multimedia | 4 |
| 2024 | Computational measurement of perceived pointiness from pronunciationabstractAbstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications. Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 5 |
| 2024 | Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 5 |
| 2023 | Discovering Phonesthemic Clusters in Readings of Kanji Characters toward Exploring Phonestheme in Japanese
Akira Yoshida, Chihaya Matsuhira, Hirotaka Kato, Takatsugu Hirayama, Takahiro Komamizu, Ichiro Ide |
PACLIC | 4 |
| 2023 | Implicit Interaction with an Autonomous Personal Mobility Vehicle: Relations of Pedestrians' Gaze Behavior with Situation Awareness and Perceived RisksabstractInteractions between pedestrians and autonomous personal mobility vehicle (APMV) will increase with the popularity of autonomous driving systems. However, when the APMVs are applied in a mixed traffic environment after manual driving PMV (MPMV) have been popular, pedestrians may feel unsafe in the interactions when they are uncertain about the driving intention of the APMV. This study seeks to find a surrogate measure for pedestrians’ understanding of driving intention and perceived safety during the interaction with an APMV. We conducted an experiment to measure the gaze duration and subjective evaluations of the participants when they interacted with a PMV in manual and autonomous driving modes. Pedestrians fixed their gaze at the APMV longer when they did not accurately understand the driving intention than when they understood it. Furthermore, the pedestrians perceived danger when they did not clearly understand the driving intention of the APMV. Besides, these factors were different when pedestrians interact with an MPMV and an APMV. Hailong Liu 0001, Takatsugu Hirayama, Luis Yoichi Morales Saiki, Hiroshi Murase |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | Detection of Localization Failures Using Markov Random Fields With Fully Connected Latent Variables for Safe LiDAR-Based Automated DrivingabstractMost of the recent automated driving systems assume the accurate functioning of localization. Unanticipated errors cause localization failures and result in failures in automated driving. An exact localization failure detection is necessary to ensure safety in automated driving; however, detection of the localization failures is challenging because sensor measurement is assumed to be independent of each other in the localization process. Owing to the assumption, the entire relation of the sensor measurement is ignored. Consequently, it is difficult to recognize the misalignment between the sensor measurement and the map when partial sensor measurement overlaps with the map. This paper proposes a method for the detection of localization failures using Markov random fields with fully connected latent variables. The full connection enables to take the entire relation into account and contributes to the exact misalignment recognition. Additionally, this paper presents localization failure probability calculation and efficient distance field representation methods. We evaluate the proposed method using two types of datasets. The first dataset is the SemanticKITTI dataset, whereby four methods are compared with the proposed method. The comparison results reveal that the proposed method achieves the most accurate failure detection. The second dataset is created based on log data acquired from the demonstrations that we conducted in Japanese public roads. The dataset includes several localization failure scenes. We apply the failure detection methods to the dataset and confirm that the proposed method achieves exact and immediate failure detection. Naoki Akai, Yasuhiro Akagi, Takatsugu Hirayama, Takayuki Morikawa, Hiroshi Murase |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Persistent Homology in LiDAR-Based Ego-Vehicle LocalizationabstractRecently, various applications leveraging topological data analysis, in particular, persistent homology (PH), have been presented in many fields since PH provides a novel point cloud analysis method. In this work, we apply PH to LiDAR-based ego-vehicle localization applications. PH can extract translation and rotation invariant features from a point cloud. These features do not maintain local information of the point cloud, such as the edges and lines; however, they can abstract the global structure of the point cloud. A persistence image (PI) vectorizes the features and allows us to obtain fixed-size vectors despite the sizes of the source point clouds being different. Additionally, the size of the PI is not large even though the source point cloud is extremely big. We consider that these advantages are effective to loop closure detection, place categorization, and end-to-end global localization applications. Results reveal that it is difficult to improve the localization accuracy by simply applying PH owing to the basic concept of the topology that does not focus on exact shapes of the geometry. Therefore, we discuss how the advantages of PH can be utilized for the localization. Naoki Akai, Takatsugu Hirayama, Hiroshi Murase |
IV | 2 |
| 2021 | Importance of Instruction for Pedestrian-Automated Driving Vehicle Interaction with an External Human Machine Interface: Effects on Pedestrians' Situation Awareness, Trust, Perceived Risks and Decision MakingabstractCompared to a manual driving vehicle (MV), an automated driving vehicle lacks a way to communicate with the pedestrian through the driver when it interacts with the pedestrian because the driver usually does not participate in driving tasks. Thus, an external human machine interface (eHMI) can be viewed as a novel explicit communication method for providing driving intentions of an automated driving vehicle (AV) to pedestrians when they need to negotiate in an interaction, e.g., an encountering scene. However, the eHMI may not guarantee that the pedestrians will fully recognize the intention of the AV. In this paper, we propose that the instruction of the eHMI's rationale can help pedestrians correctly understand the driving intentions and predict the behavior of the AV, and thus their subjective feelings (i. e., dangerous feeling, trust in the AV, and feeling of relief) and decision-making are also improved. The results of an interaction experiment in a road-crossing scene indicate that the participants were more difficult to be aware of the situation when they encountered an AV w/o eHMI compared to when they encountered an MV; further, the participants' subjective feelings and hesitation in decision-making also deteriorated significantly. When the eHMI was used in the AV, the situational awareness, subjective feelings and decision-making of the participants regarding the AV w/ eHMI were improved. After the instruction, it was easier for the participants to understand the driving intention and predict driving behavior of the AV w/ eHMI. Further, the subjective feelings and the hesitation related to decision-making were improved and reached the same standards as that for the MV. Hailong Liu 0001, Takatsugu Hirayama, Masaya Watanabe |
IV | 2 |
| 2021 | Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 5 |
| 2021 | Experimental stability analysis of neural networks in classification problems with confidence sets for persistence diagrams
Naoki Akai, Takatsugu Hirayama, Hiroshi Murase |
Neural Networks | 2 |
| 2020 | Hybrid Localization using Model- and Learning-Based Methods: Fusion of Monte Carlo and E2E Localizations via Importance SamplingabstractThis paper proposes a hybrid localization method that fuses Monte Carlo localization (MCL) and convolutional neural network (CNN)-based end-to-end (E2E) localization. MCL is based on particle filter and requires proposal distributions to sample the particles. The proposal distribution is generally predicted using a motion model. However, because the motion model cannot handle unanticipated errors, the predicted distribution is sometimes inaccurate. The use of other ideal proposal distributions, such as the measurement model, can improve robustness against such unanticipated errors. This technique is called importance sampling (IS). However, it is difficult to sample the particles from such ideal distributions because they are not represented in the closed form. Recent works have proved that CNNs with dropout layers represent the posterior distributions over their outputs conditioned on the inputs and the CNN predictions are equivalent to sampling the outputs from the posterior. Therefore, the proposed method utilizes a CNN to sample the particles and fuses them with MCL via IS. Consequently, the advantages of both MCL and E2E localization can be simultaneously leveraged while preventing their disadvantages. Experiments demonstrate that the proposed method can smoothly estimate the robot pose, similar to the model-based method, and quickly re-localize it from the failures, similar to the learning-based method. Naoki Akai, Takatsugu Hirayama, Hiroshi Murase |
ICRA | 2 |
| 2020 | 3D Monte Carlo Localization with Efficient Distance Field Representation for Automated Driving in Dynamic EnvironmentsabstractThis paper presents a LiDAR-based 3D Monte Carlo localization (MCL) with an efficient distance field (DF) representation method. To implement 3D MCL, high computing capacity is required because the likelihood of many pose candidates, i.e., particles, must be calculated in real time by comparing sensor measurements and a map. Additionally, a large-scale map is needed for allocation to embedded computers since autonomous vehicles are required to navigate wide areas. These make it difficult for 3D MCL implementation. This paper first presents an efficient DF representation method while considering the 3D LiDAR-based localization characteristics. Because each DF voxel has the closest distance from occupied voxels, swift comparison of the sensor measurements and map can be achieved. Consequently, 3D MCL using the likelihood field model (LFM) can be executed in real time. Furthermore, this paper presents a method for improving the localization robustness to environmental changes without increasing memory and computational cost from that of the LFM-based MCL. Through experiments using the SemanticKITTI dataset, we show that the presented method can efficiently and robustly work in dynamic environments. Naoki Akai, Takatsugu Hirayama, Hiroshi Murase |
IV | 2 |
| 2020 | Automatic Interaction Detection Between Vehicles and Vulnerable Road Users During Turning at an IntersectionabstractInteraction detection between vehicles and vulnerable road users (e.g. pedestrians and cyclists) is important for e.g. safety control and autonomous driving. However, there are many challenges for automatically detecting interactions, such as the ambiguity of defining when interaction is required in dynamic traffic activities among different road users and the lack of labeled data for training a machine learning detector. To overcome the challenges, we introduce a way to define whether or not interaction is required in various traffic scenes and create a large real-world dataset from a very challenging intersection. A sequence-to-sequence method that uses the object information and motion information of the traffic scenes extracted by a state-of-the-art object detector and from optical flow, respectively, is proposed for automatic interaction detection. The proposed method generates a probability of interaction at each short interval (<; 0.1 s) that represents the changing of interaction along a sequence. We obtain a baseline model that differentiates no interaction from interaction on the basis of the location and road user type from the detected object information. Compared with the baseline model, the empirical results of the proposed method demonstrate very accurate predictions for vehicle turning sequences with varying length. Hao Cheng 0008, Hailong Liu 0001, Fumito Shinmura, Naoki Akai, Hiroshi Murase, Takatsugu Hirayama |
IV | 6 |
| 2020 | Imageability Estimation using Visual and Language FeaturesabstractImageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability. Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
ICMR | 5 |
| 2020 | Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 4 |
| 2020 | More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 2 |
| 2020 | Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 5 |
| 2019 | Driving Behavior Modeling Based on Hidden Markov Models with Driver's Eye-Gaze Measurement and Ego-Vehicle LocalizationabstractThis paper presents a comparison of driving behavior modeling methods based on hidden Markov models (HMMs) with driver's eye-gaze measurement and ego-vehicle localization. Original HMMs are sometimes insufficient to model real-world scenarios. To overcome these limitations, extended HMMs have been proposed, e.g., autoregressive input-output HMMs (AIOHMMs). This paper first details AIOHMMs and presents ways to use them for driving behavior modeling. We compare the performance for behavior modeling and maneuver discrimination for six types of HMMs. The driving data for this work was gathered in our university campus with a car-like vehicle. Experimental results suggest that the hidden states can properly represent the average of the driving actions when the driving behaviors are accurately modeled by the HMMs. It is also suggested that surrounding and past information can be used to flexibly model the relationship between driving actions and related information. Naoki Akai, Takatsugu Hirayama, Luis Yoichi Morales Saiki, Yasuhiro Akagi, Hailong Liu 0001, Hiroshi Murase |
IV | 2 |
| 2019 | Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 4 |
| 2018 | Gaze-Inspired Learning for Estimating the Attractiveness of a Food PhotoabstractThe number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them. Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ISM | 2 |
| 2016 | A classification method of cooking operations based on eye movement patternsabstractWe are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables. Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ETRA | 2 |
| 2016 | Skilled gaze behavior extraction based on dependency analysis of gaze patterns on video scenesabstractThe eye gaze behavior of individuals changes depending on their knowledge and experience of the event occurring in their field of view. In past studies, researchers formulated a hypothesis concerning this dependency on a specific scene and then analyzed the gaze behavior of viewers observing the scene. We depart from this hypothesis-testing paradigm. In this paper, we propose a data-mining framework for extracting skilled gaze behaviors of experts while watching a video based on a comprehensive comparison of viewers in terms of the dependency of their gaze patterns on video scenes. To quantitatively analyze the changes in the gaze behavior of experts according to the events in the scene, video and eye movement sequences are classified into video scenes and gaze patterns, respectively, by using an unsupervised clustering method focusing on short-time dynamics. Then, we analyze the dependency based on the distinctiveness and occurrence frequency of gaze patterns for each video scene. Atsushi Iwatsuki, Takatsugu Hirayama, Junya Morita, Kenji Mase |
ETRA | 2 |
| 2016 | Model-based Reminiscence: Guiding Mental Time Travel by Cognitive ModelingabstractThis paper proposes an approach to elderly mental care called model-based reminiscence, which utilizes cognitive modeling to guide a user's mental time travel. In this approach, a personalized cognitive model is constructed by implementing a user's lifelog (a photo library) in the ACT-R cognitive architecture. The constructed model retrieves photos based on human memory characteristics such as learning, forgetting, inhibition, and noise. These memory characteristics are regulated with parameter values corresponding to cognitive and emotional health. The authors assumed that a user's mental health could be assessed from their reactions to photo sequences retrieved by models with various parameter settings. The authors also assumed that it would be possible to motivate a user by guiding their memory recall with photo sequences generated from a healthy optimal state model. A simulation study indicates the potential of this approach presenting a variety of model behaviors corresponding cognitive / emotional states. Junya Morita, Takatsugu Hirayama, Kenji Mase, Kazunori Yamada |
HAI | 2 |
| 2016 | Multimodal biofeedback system integrating low-cost easy sensing devicesabstractWe built a multimodal biofeedback system integrated with low-cost sensors and applied the system to support meditation training utilizing multimodal biofeedback to the user. Our meditation support system employs Electroencephalogram, heart rate variability and eye tracking. The first two biosignals are employed to assess the mental stress during meditation, and eye tracking is equipped to detect an interval where the user engages in meditation by monitoring the open/closed states of the eyes. Wataru Hashiguchi, Junya Morita, Takatsugu Hirayama, Kenji Mase, Kazunori Yamada, Mayu Yokoya |
ICMI | 3 |
| 2016 | Personal Multi-view Viewpoint Recommendation based on Trajectory Distribution of the Viewing TargetabstractMulti-camera videos with abundant information and high flexibility are expected to be useful in a wide range of applications, such as surveillance systems, web lecture broadcasting, concerts and sports viewing, etc. Viewers can enjoy a high-presence viewing experience of their own choosing by means of virtual camera switching and controlling viewing interfaces. However, some viewers may feel annoyed by continual manual viewpoint selection, especially when the number of selectable viewpoints is relatively large. In order to solve this issue, we propose an automatic viewpoint-recommending method designed especially for soccer games. This method focuses on a viewer's personal preference for viewpoint-selection, instead of common and professional editing rules. We assume that the different trajectory distributions cause a difference in the viewpoint selection according to personal preference. We therefore analyze the relationship between the viewer's personal viewpoint selecting tendency and the spatio-temporal game context. We compare methods based on a Gaussian mixture model, a general histogram+SVM and bag-of-words+SVM to seek the best representation for this relationship. The performance of the proposed methods are verified by assessing the degree of similarity between the recommended viewpoints and the viewers' edited records. Kensho Hara, Yu Enokibori, Takatsugu Hirayama, Kenji Mase |
ACM Multimedia | 4 |
| 2015 | Cognitive Modeling of Life Story: Reconstructing Our Memories from a Photo Library
Junya Morita, Takatsugu Hirayama, Kenji Mase, Kazunori Yamada |
CogSci | 2 |
| 2015 | Analyzing driver gaze behavior and consistency of decision making during automated drivingabstractWe investigate a possible method for detecting a driver's negative adaptation to an automated driving system by analyzing consistency of driver decision making and driver gaze behavior during automated driving. We focus on an automated driving system equivalent to Level 2 automation per the NHTSA's definition. At this level of automation, drivers must be ready to take control of the vehicle in critical situations by monitoring the driving environment and vehicle behavior. Since drivers are not required to operate the pedals or steering wheel during automated driving, a driver's negative adaptation to an automated system needs to be detected from behavior other than vehicle operation. In this study, we focus on driver gaze behavior. We conduct a simulator study to compare the gaze behavior of fifteen drivers during conventional and automated driving. We also analyze the consistency of driver decision making when changing lanes during conventional and automated driving. Experimental results show that drivers who pay less attention to the road ahead during automated driving tend to be less sensitive to risk factors in the surrounding environment and also tend to make inconsistent lane change decisions during automated driving. Chiyomi Miyajima, Suguru Yamazaki, Takashi Bando, Kentarou Hitomi, Hitoshi Terai, Hiroyuki Okuda, Takatsugu Hirayama, Masumi Egawa, Tatsuya Suzuki 0001, Kazuya Takeda |
Intelligent Vehicles Symposium | 7 |
| 2014 | Analysis of gaze behavior while using a multi-viewpoint video viewerabstractHumans see things from various viewpoints but nobody attempts to see anything from every viewpoint owing to physical limitations and the great effort required. Intelligent interfaces for viewing multi-viewpoint videos may effectively remove these limitations and open up a new visual world to mankind. We have developed a multi-viewpoint video viewer that incorporates target-centered viewpoint switching. The viewer stabilizes an object at the center of the display field, which helps to focus the user's gaze on the target. We conducted a user study to analyze user behavior, especially eye movement, while watching a multi-viewpoint video on the viewer. Statistical analyses of the results indicated that the target-centered viewpoint switching encouraged the users to gaze at the center of the display where the target was located during the viewing. We believe that these are useful findings that pave the way for the design of even more intelligent viewers. Takatsugu Hirayama, Takafumi Marutani, Sidney S. Fels, Kenji Mase |
ETRA | 1 |
| 2014 | Constructing a Non-task-oriented Dialogue Agent using Statistical Response Method and GamificationabstractThis paper provides a novel method for building non-task-oriented dialogue agents such as chatbots. The
dialogue agent constructed using our method automatically selects a suitable utterance depending on a context
from a set of candidate utterances prepared in advance. To realize automatic utterance selection, we rank the
candidate utterances in order of suitability by application of a machine learning algorithm. We employed both
right and wrong dialogue data to learn relative suitability to rank the utterances. Additionally, we provide
a low-cost and quality-assured learning data acquisition environment using crowdsourcing and gamification.
The results of an experiment using learning data obtained via the environment demonstrate that the appropriate
utterance is ranked on the top in 82.6% of cases and within the top 3 at 95.0% of cases. Results show that
using context information that is not used in most existing agents is necessary for appropriate responses. Michimasa Inaba, Naoyuki Iwata, Fujio Toriumi, Takatsugu Hirayama, Yu Enokibori, Kenichi Takahashi, Kenji Mase |
ICAART (1) | 4 |
| 2014 | Trend-sensitive hough forests for action detectionabstractA Hough transform-based method for action detection can achieve robustness to occlusions because the method casts votes for action classes and spatio-temporal action positions based on the visible local features of partially occluded actions. However, each local feature is prone to a false vote. This paper focuses on the trend of past votes to curb the influence of false votes by extending conventional Hough forests to sensing that trend. Our proposed method, called trendsensitive Hough forests, learns a voting trend model that can be used to discriminate between correct and false votes and calculate the confidence of them. We experimentally confirmed that it outperformed action detection accuracy of conventional Hough forests. Kensho Hara, Takatsugu Hirayama, Kenji Mase |
ICIP | 2 |
| 2014 | Context-Dependent Viewpoint Sequence Recommendation System for Multi-view VideoabstractMulti-view videos shot using multiple cameras are highly interested due to their considerable flexibility in enhancing the quality of our daily viewing experience, especially for large-scale events. However, the increase in the number of cameras burdens even experts on suitable viewpoint selection. Therefore, we propose in this paper an automatic viewpoint sequence recommendation system to support multi-view viewpoint selecting with a soccer game example. Unlike existing methods, our proposed system focuses on context-dependency using viewpoint evaluation and transition processes by two types of agents: a camera agent and a producer agent. The camera agent evaluates the view quality based on scene context such as positions of ball and players in given production context such as camera position and user's preference. The producer agent selects the optimal set of viewpoints by taking account of the view quality and the production objectives. The context-dependent optimization has been performed to generate variable viewing patterns which are adequate to various scene and production contexts. Sequences generated by the system and the human selection were experimentally compared to confirm the effectiveness of our proposed system. Our recommendation system has the potential to satisfy both common and personal viewing preferences for sports games. Yuki Muramatsu, Takatsugu Hirayama, Kenji Mase |
ISM | 3 |
| 2010 | Gaze Probing: Event-Based Estimation of Objects Being Focused OnabstractWe propose a novel method to estimate the object that a user is focusing on by using the synchronization between the movements of objects and a user's eyes as a cue. We first design an event as a characteristic motion pattern, and we then embed it within the movement of each object. Since the user's ocular reactions to these events are easily detected using a passive camera-based eye tracker, we can successfully estimate the object that the user is focusing on as the one whose movement is most synchronized with the user's eye reaction. Experimental results obtained from the application of this system to dynamic content (consisting of scrolling images) demonstrate the effectiveness of the proposed method over existing methods. Ryo Yonetani, Hiroaki Kawashima, Takatsugu Hirayama, Takashi Matsuyama |
ICPR | 3 |
| 2008 | Person-independent face tracking based on dynamic AAM selectionabstractWe have developed a high-precision method that selects an appropriate model of a video image in order to track an unknown face in front of a large display. Currently, Active Appearance Models (AAMs) are used to track non-rigid objects, such as a faces, because the models efficiently learn the correlation between shape and texture. The problem with an AAM is that when it tracks an unknown face, excessive training data increases tracking errors because there is an intermediate model size beyond which the reduction in fitting performance outweighs the gains from any improved representational power of the model. To increases the accuracy with which an unknown face is tracked, we built clustered models from training datasets and select a cluster that includes a face which is similar to the unknown face. Our method of clustering and cluster selecting is based on the Mutual Subspace Method (MSM). We demonstrated the effectiveness of our method by using the leave-one-out cross-validation. Akihiro Kobayashi, Junji Satake, Takatsugu Hirayama, Hiroaki Kawashima, Takashi Matsuyama |
FG | 3 |