EDBT 2026 Demo / reviewers in the wild / expert
Shigeo Morishima
dblp:60/4254
· DBLP profile ↗
101ranked-venue papers
11as first author
28since 2021 · last 2025
0000-0001-8859-6539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 22 · 1 first-author · 11 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WanderGuide: Indoor Map-less Robotic Guide for Exploration by Blind PeopleabstractBlind people have limited opportunities to explore an environment based on their interests.While existing navigation systems could provide them with surrounding information while navigating, they have limited scalability as they require preparing prebuilt maps.Thus, to develop a map-less robot that assists blind people in exploring, we first conducted a study with ten blind participants at a shopping mall and science museum to investigate the requirements of the system, which revealed the need for three levels of detail to describe the surroundings based on users' preferences.Then, we developed WanderGuide, with functionalities that allow users to adjust the level of detail in descriptions and verbally interact with the system to ask questions about the environment or to go to points of interest.The study with five blind participants revealed that WanderGuide could provide blind people with the enjoyable experience of wandering around without a specific destination in their minds. Masaki Kuribayashi, Kohei Uehara, Allan Wang, Shigeo Morishima, Chieko Asakawa |
CHI | 4 |
| 2025 | Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions
Yusuke Miura, Chi-Lan Yang, Masaki Kuribayashi, Keigo Matsumoto, Hideaki Kuzuoka, Shigeo Morishima |
CHI | 6 |
| 2025 | Viewpoint-Dependent 3D Visual Grounding for Mobile Robotsabstract3D visual grounding is the task of identifying objects in spatial environments based on textual descriptions, enabling natural language interactions between humans and robots. However, existing studies overlook viewpoint-dependent texts expressions, such as "the chair to your right", despite their frequent use in human instructions. In this paper, we introduce a novel problem setting focused on viewpoint-dependent texts and present a new dataset that incorporates the robot’s viewpoint. We conducted three experiments to analyze the dataset’s difficulty and characteristics by comparing existing models and newly designed models that take viewpoint as an additional input. Our results indicate that considering viewpoint is crucial for the object selection process in our task. In addition, we also found that the difficulty of the task varies depending on how the object is described in the text. Shogo Iwakata, Ryosuke Oshima, Hideki Tsunashima, Hirokatsu Kataoka, Shigeo Morishima |
ICIP | 6 |
| 2025 | Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
Shunsuke Mitsumori, Sara Kashiwagi, Keitaro Tanaka, Shigeo Morishima |
INTERSPEECH | 4 |
| 2025 | Training Onset-and-Offset-Aware Sound Event Detection on a Heterogeneous Dataset via Probabilistic Sequential Modeling
Tomoya Yoshinaga, Yoshiaki Bando, Keitaro Tanaka, Keisuke Imoto, Masaki Onishi, Shigeo Morishima |
INTERSPEECH | 6 |
| 2025 | SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and FingeringabstractAutomatically generating realistic musical performance motion can greatly enhance digital media production, often involving collaboration between professionals and musicians. However, capturing the intricate body, hand, and finger movements required for accurate musical performances is challenging. Existing methods often fall short due to the complex mapping between audio and motion, typically requiring additional inputs like scores or MIDI data. In this work, we present SyncViolinist, a multi-stage end-to-end framework that generates synchronized violin performance motion solely from audio input. Our method overcomes the challenge of capturing both global and finegrained performance features through two key modules: a bowing/fingering module and a motion generation module. The bowing/fingering module extracts detailed playing information from the audio, which the motion generation module uses to create precise, coordinated body motions reflecting the temporal granularity and nature of the violin performance. We demonstrate the effectiveness of SyncViolinist with significantly improved qualitative and quantitative results from unseen violin performance audio, outperforming state-of-the-art methods. Extensive subjective evaluations involving professional violinists further validate our approach. The code and dataset are available at https://github.com/Kakanat/SyncViolinist. Hiroki Nishizawa, Keitaro Tanaka, Asuka Hirata, Shugo Yamaguchi, Masatoshi Hamanaka, Shigeo Morishima |
WACV | 7 |
| 2025 | Geometric visual fusion graph neural networks for multi-person human-object interaction recognition in videosabstractHuman-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture appearance context, while geometric features provide structural patterns. Effectively fusing these multimodal features without compromising their unique characteristics remains challenging. We observe that establishing robust, entity-specific representations before modeling interactions helps preserve the strengths of each modality. Therefore, we hypothesize that a bottom-up approach is crucial for effective multimodal fusion. Following this insight, we propose the Geometric Visual Fusion Graph Neural Network (GeoVis-GNN), which uses dual-attention feature fusion combined with interdependent entity graph learning. It progressively builds from entity-specific representations toward high-level interaction understanding. To advance HOI recognition to real-world scenarios, we introduce the Concurrent Partial Interaction Dataset (MPHOI-120). It captures dynamic multi-person interactions involving concurrent actions and partial engagement. This dataset helps address challenges like complex human-object dynamics and mutual occlusions. Extensive experiments demonstrate the effectiveness of our method across various HOI scenarios. These scenarios include two-person interactions, single-person activities, bimanual manipulations, and complex concurrent partial interactions. Our method achieves state-of-the-art performance. Tanqiu Qiao, Ruochen Li 0002, Frederick W. B. Li, Yoshiki Kubotani, Shigeo Morishima, Hubert P. H. Shum |
Expert Syst. Appl. | 5 |
| 2025 | Onset-and-Offset-Aware Sound Event Detection via Differentiable Frame-to-Event MappingabstractThis paper presents a sound event detection (SED) method that handles sound event boundaries in a statistically principled manner. A typical approach to SED is to train a deep neural network (DNN) in a supervised manner such that the model predicts frame-wise event activities. Since the predicted activities often contain fine insertion and deletion errors due to their temporal fluctuations, post-processing has been applied to obtain more accurate onset and offset boundaries. Existing post-processing methods are, however, non-differentiable and prohibit end-to-end (E2E) training. In this paper, we propose an E2E detection method based on a probabilistic formulation of sound event sequences called a hidden semi-Markov model (HSMM). The HSMM is utilized to transform frame-wise features predicted by a DNN into posterior probabilities of sound events represented by their class labels and temporal boundaries. We jointly train the DNN and HSMM in a supervised E2E manner by maximizing the event-wise posterior probabilities of the HSMM. This objective is a differentiable function thanks to the forward-backward algorithm of the HSMM. Experimental results with real recordings show that our method outperforms baseline systems with standard post-processing methods. Tomoya Yoshinaga, Keitaro Tanaka, Yoshiaki Bando, Keisuke Imoto, Shigeo Morishima |
IEEE Signal Process. Lett. | 5 |
| 2024 | Monte Carlo Path Tracing and Statistical Event Detection for Event Camera SimulationabstractThis paper presents a novel event camera simulation system fully based on physically based Monte Carlo path tracing with adaptive path sampling. The adaptive sampling performed in the proposed method is based on a statistical technique, hypothesis testing for the hypothesis whether the difference of logarithmic luminances at two distant periods is significantly larger than a predefined event threshold. To this end, our rendering system collects logarithmic luminances rather than raw luminance in contrast to the conventional rendering system imitating conventional RGB cameras. Then, based on the central limit theorem, we reasonably assume that the distribution of the population mean of logarithmic luminance can be modeled as a normal distribution, allowing us to model the distribution of the difference of logarithmic luminance as a normal distribution. Then, using Student's t-test, we can test the hypothesis and determine whether to discard the null hypothesis for event non-occurrence. When we sample a sufficiently large number of path samples to satisfy the central limit theorem and obtain a clean set of events, our method achieves significant speed up compared to a simple approach of sampling paths uniformly at every pixel. To our knowledge, we are the first to simulate the behavior of event cameras in a physically accurate manner using an adaptive sampling technique in Monte Carlo path tracing, and we believe this study will contribute to the development of computer vision applications using event cameras. Yuichiro Manabe, Tatsuya Yatagawa, Shigeo Morishima, Hiroyuki Kubo |
ICCP | 3 |
| 2024 | The Gap in the Strategy of Recovering Task Failure between GPT-4V and Humans in a Visual DialogueabstractGoal-oriented dialogue systems interact with humans to accomplish specific tasks.However, sometimes these systems fail to establish a common ground with users, leading to task failures.In such cases, it is crucial not to just end with failure but to correct and recover the dialogue to turn it into a success for building a robust goal-oriented dialogue system.Effective recovery from task failures in a goal-oriented dialogue involves not only successful recovery but also accurately understanding the situation of the failed task to minimize unnecessary interactions and avoid frustrating the user.In this study, we analyze the capabilities of GPT-4V in recovering failure tasks by comparing its performance with that of humans using Guess What?! Game.The results show that GPT-4V employs less efficient recovery strategies, such as asking additional unnecessary questions, than humans.We also found that while humans can occasionally ask questions that doubt the accuracy of the interlocutor's answer during task recovery, GPT-4V lacks this capability. Ryosuke Oshima, Seitaro Shinagawa, Shigeo Morishima |
SIGDIAL | 3 |
| 2024 | ChitChatGuide: Conversational Interaction Using Large Language Models for Assisting People with Visual Impairments to Explore a Shopping MallabstractTo enable people with visual impairments (PVI) to explore shopping malls, it is important to provide information for selecting destinations and obtaining information based on the individual's interests. We achieved this through conversational interaction by integrating a large language model (LLM) with a navigation system. ChitChatGuide allows users to plan a tour through contextual conversations, receive personalized descriptions of surroundings based on transit time, and make inquiries during navigation. We conducted a study in a shopping mall with 11 PVI, and the results reveal that the system allowed them to explore the facility with increased enjoyment. The LLM-based conversational interaction, by understanding vague and context-based questions, enabled the participants to explore unfamiliar environments effectively. The personalized and in-situ information generated by the LLM was both useful and enjoyable. Considering the limitations we identified, we discuss the criteria for integrating LLMs into navigation systems to enhance the exploration experiences of PVI. Yuka Kaniwa, Masaki Kuribayashi, Seita Kayukawa, Daisuke Sato 0001, Hironobu Takagi, Chieko Asakawa, Shigeo Morishima |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2024 | Snap&Nav: Smartphone-based Indoor Navigation System For Blind People via Floor Map Analysis and Intersection DetectionabstractWe present Snap&Nav, a navigation system for blind people in unfamiliar buildings, without prebuilt digital maps. Instead, the system utilizes the floor map as its primary information source for route guidance. The system requires a sighted assistant to capture an image of the floor map, which is analyzed to create a node map containing intersections, destinations, and current positions on the floor. The system provides turn-by-turn navigation instructions while tracking users' positions on the node map by detecting intersections. Additionally, the system estimates the scale difference of the node map to provide distance information. Our system was validated through two user studies with 20 sighted and 12 blind participants. Results showed that sighted participants processed floor map images without being accustomed to the system, while blind participants navigated with increased confidence and lower cognitive load compared to the condition using only cane, appreciating the system's potential for use in various buildings. Masaya Kubota, Masaki Kuribayashi, Seita Kayukawa, Hironobu Takagi, Chieko Asakawa, Shigeo Morishima |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2023 | Enhancing Blind Visitor's Autonomy in a Science Museum Using an Autonomous Navigation RobotabstractEnabling blind visitors to explore museum floors while feeling the facility’s atmosphere and increasing their autonomy and enjoyment are imperative for giving them a high-quality museum experience. We designed a science museum exploration system for blind visitors using an autonomous navigation robot. Blind users can control the robot to navigate them toward desired exhibits while playing short audio descriptions along the route. They can also browse detailed explanations on their smartphones and call museum staff if interactive support is needed. Our real-world user study at a science museum during its opening hour revealed that blind participants could explore the museum safely and independently at their own pace. The study also showed that the sighted visitors who saw the participants walking with the robot accepted the assistive robot well. We finally conducted focus group sessions with the blind participants and discussed further requirements toward a more independent museum experience. Seita Kayukawa, Daisuke Sato 0001, Masayuki Murata 0002, Tatsuya Ishihara, Hironobu Takagi, Shigeo Morishima, Chieko Asakawa |
CHI | 6 |
| 2023 | PathFinder: Designing a Map-less Navigation System for Blind People in Unfamiliar BuildingsabstractIndoor navigation systems with prebuilt maps have shown great potential in navigating blind people even in unfamiliar buildings. However, blind people cannot always benefit from them in every building, as prebuilt maps are expensive to build. This paper explores a map-less navigation system for blind people to reach destinations in unfamiliar buildings, which is implemented on a robot. We first conducted a participatory design with five blind people, which revealed that intersections and signs are the most relevant information in unfamiliar buildings. Then, we prototyped PathFinder, a navigation system that allows blind people to determine their way by detecting and conveying information about intersections and signs. Through a participatory study, we improved the interface of PathFinder, such as the feedback for conveying the detection results. Finally, a study with seven blind participants validated that PathFinder could assist users in navigating unfamiliar buildings with increased confidence compared to their regular aid. Masaki Kuribayashi, Tatsuya Ishihara, Daisuke Sato 0001, Jayakorn Vongkulbhisal, Karnik Ram, Seita Kayukawa, Hironobu Takagi, Shigeo Morishima, Chieko Asakawa |
CHI | 8 |
| 2023 | Scapegoat Generation for Privacy Protection from DeepfakeabstractTo protect privacy and prevent malicious use of deepfake, current studies propose methods that interfere with the generation process, such as detection and destruction approaches. However, these methods suffer from sub-optimal generalization performance to unseen models and add undesirable noise to the original image. To address these problems, we propose a new problem formulation for deepfake prevention: generating a "scapegoat image" by modifying the style of the original input in a way that is recognizable as an avatar by the user, but impossible to reconstruct the real face. Even in the case of malicious deepfake, the privacy of the users is still protected. To achieve this, we introduce an optimization-based editing method that utilizes GAN inversion to discourage deepfake models from generating similar scapegoats. We validate the effectiveness of our proposed method through quantitative and user studies. Gido Kato, Yoshihiro Fukuhara, Mariko Isogawa, Hideki Tsunashima, Hirokatsu Kataoka, Shigeo Morishima |
ICIP | 6 |
| 2023 | Event-Based Camera Simulation Using Monte Carlo Path Tracing with Adaptive DenoisingabstractThis paper presents an algorithm to obtain an event-based video from noisy frames given by physics-based Monte Carlo path tracing over a synthetic 3D scene. Given the nature of dynamic vision sensor (DVS), rendering event-based video can be viewed as a process of detecting the changes from noisy brightness values. We extend a denoising method based on a weighted local regression (WLR) to detect the brightness changes rather than applying denoising to every pixel. Specifically, we derive a threshold to determine the likelihood of event occurrence and reduce the number of times to perform the regression. Our method is robust to noisy video frames obtained from a few path-traced samples. Despite its efficiency, our method performs comparably to or even better than an approach that exhaustively denoises every frame. Visit our project page for more information: https: //github.com/0V/ESIM-AD.git. Yuta Tsuji, Tatsuya Yatagawa, Hiroyuki Kubo, Shigeo Morishima |
ICIP | 4 |
| 2023 | Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric LearningabstractThis paper presents a novel metric learning approach to address the performance gap between normal and silent speech in visual speech recognition (VSR).The difference in lip movements between the two poses a challenge for existing VSR models, which exhibit degraded accuracy when applied to silent speech.To solve this issue and tackle the scarcity of training data for silent speech, we propose to leverage the shared literal content between normal and silent speech and present a metric learning approach based on visemes.Specifically, we aim to map the input of two speech types close to each other in a latent space if they have similar viseme representations.By minimizing the Kullback-Leibler divergence of the predicted viseme probability distributions between and within the two speech types, our model effectively learns and predicts viseme identities.Our evaluation demonstrates that our method improves the accuracy of silent VSR, even when limited training data is available. Sara Kashiwagi, Keitaro Tanaka, Shigeo Morishima |
INTERSPEECH | 4 |
| 2023 | Enhancing Perception and Immersion in Pre-Captured Environments through Learning-Based Eye Height AdaptationabstractPre-captured immersive environments using omnidirectional cameras provide a wide range of virtual reality applications. Previous research has shown that manipulating the eye height in egocentric virtual environments can significantly affect distance perception and immersion. However, the influence of eye height in pre-captured real environments has received less attention due to the difficulty of altering the perspective after finishing the capture process. To explore this influence, we first propose a pilot study that captures real environments with multiple eye heights and asks participants to judge the egocentric distances and immersion. If a significant influence is confirmed, an effective image-based approach to adapt pre-captured real-world environments to the user’s eye height would be desirable. Motivated by the study, we propose a learning-based approach for synthesizing novel views for omnidirectional images with altered eye heights. This approach employs a multitask architecture that learns depth and semantic segmentation in two formats, and generates high-quality depth and semantic segmentation to facilitate the inpainting stage. With the improved omnidirectional-aware layered depth image, our approach synthesizes natural and realistic visuals for eye height adaptation. Quantitative and qualitative evaluation shows favorable results against state-of-the-art methods, and an extensive user study verifies improved perception and immersion for pre-captured real-world environments. Hubert P. H. Shum, Shigeo Morishima |
ISMAR | 3 |
| 2023 | A Conservative Semi-Implicit Scheme for Shallow Water EquationsabstractThis paper presents a new mass and (local) momentum conserving scheme for shallow water equations that are discretized in conservation form. In comparison to conventional techniques, such as the semi-Lagrangian scheme and its conservative variant, our approach offers noticeably improved visual animation of wave dynamics. Haruka Hirae, Shigeo Morishima, Ryoichi Ando |
SCA | 2 |
| 2022 | Geometric Features Informed Multi-person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Qianhui Men, Frederick W. B. Li, Yoshiki Kubotani, Shigeo Morishima, Hubert P. H. Shum |
ECCV (4) | 5 |
| 2022 | The Sound of Bounding-BoxesabstractIn the task of audio-visual sound source separation, which leverages visual information for sound source separation, identifying objects in an image is a crucial step prior to separating the sound source. However, existing methods that assign sound on detected bounding boxes suffer from a problem that their approach heavily relies on pre-trained object detectors. Specifically, when using these existing methods, it is required to predetermine all the possible categories of objects that can produce sound and use an object detector applicable to all such categories. To tackle this problem, we propose a fully unsupervised method that learns to detect objects in an image and separate sound source simultaneously. As our method does not rely on any pre-trained detector, our method is applicable to arbitrary categories without any additional annotation. Furthermore, although being fully unsupervised, we found that our method performs comparably in separation accuracy. Takashi Oya, Shohei Iwase, Shigeo Morishima |
ICPR | 3 |
| 2022 | How Users, Facility Managers, and Bystanders Perceive and Accept a Navigation Robot for Visually Impaired People in Public BuildingsabstractAutonomous navigation robots have a considerable potential to offer a new form of mobility aid to people with visual impairments. However, to deploy such robots in public buildings, it is imperative to receive acceptance from not only robot users but also people that use the buildings and managers of those facilities. Therefore, we conducted three studies to investigate the acceptance and concerns of our prototype robot, which looks like a regular suitcase. First, an online survey revealed that people could accept the robot navigating blind users. Second, in the interviews with facility managers, they were cautious about the robot’s camera and the privacy of their customers. Finally, focus group sessions with legally blind participants who experienced the robot navigation revealed that the robot may cause trouble when it collides with those who may not be aware of the user’s blindness. Still, many participants liked the design of the robot which assimilated into the surroundings. Seita Kayukawa, Daisuke Sato 0001, Masayuki Murata 0002, Tatsuya Ishihara, Akihiro Kosugi, Hironobu Takagi, Shigeo Morishima, Chieko Asakawa |
RO-MAN | 7 |
| 2022 | 360 Depth Estimation in the Wild - the Depth360 Dataset and the SegFuse NetworkabstractSingle-view depth estimation from omnidirectional images has gained popularity with its wide range of applications such as autonomous driving and scene reconstruction. Although data-driven learning-based methods demonstrate significant potential in this field, scarce training data and ineffective 360 estimation algorithms are still two key limitations hindering accurate estimation across diverse domains. In this work, we first establish a large-scale dataset with varied settings called Depth360 to tackle the training data problem. This is achieved by exploring the use of a plenteous source of data, 360 videos from the internet, using a test-time training method that leverages unique information in each omnidirectional sequence. With novel geometric and temporal constraints, our method generates consistent and convincing depth samples to facilitate single-view estimation. We then propose an end-to-end two-branch multi-task learning network, SegFuse, that mimics the human eye to effectively learn from the dataset and estimate high-quality depth maps from diverse monocular RGB images. With a peripheral branch that uses equirectangular projection for depth estimation and a foveal branch that uses cubemap projection for semantic segmentation, our method predicts consistent global depth while maintaining sharp details at local regions. Experimental results show favorable performance against the state-of-the-art methods. Hubert P. H. Shum, Shigeo Morishima |
VR | 3 |
| 2022 | Corridor-Walker: Mobile Indoor Walking Assistance for Blind People to Avoid Obstacles and Recognize IntersectionsabstractNavigating in an indoor corridor can be challenging for blind people as they have to be aware of obstacles while also having to recognize the intersections that lead to the destination. To aid blind people in such tasks, we propose Corridor-Walker, a smartphone-based system that assists blind people to avoid obstacles and recognize intersections. The system uses a LiDAR sensor equipped with a smartphone to construct a 2D occupancy grid map of the surrounding environment. Then, the system generates an obstacle-avoiding path and detects upcoming intersections on the grid map. Finally, the system navigates the user to trace the generated path and notifies the user of each intersection's existence and the shape using vibration and audio feedback. A user study with 14 blind participants revealed that Corridor-Walker allowed participants to avoid obstacles, rely less on the wall to walk straight, and enable them to recognize intersections. Masaki Kuribayashi, Seita Kayukawa, Jayakorn Vongkulbhisal, Chieko Asakawa, Daisuke Sato 0001, Hironobu Takagi, Shigeo Morishima |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2022 | 3D car shape reconstruction from a contour sketch using GAN and lazy learningabstractAbstract 3D car models are heavily used in computer games, visual effects, and even automotive designs. As a result, producing such models with minimal labour costs is increasingly more important. To tackle the challenge, we propose a novel system to reconstruct a 3D car using a single sketch image. The system learns from a synthetic database of 3D car models and their corresponding 2D contour sketches and segmentation masks, allowing effective training with minimal data collection cost. The core of the system is a machine learning pipeline that combines the use of a generative adversarial network (GAN) and lazy learning. GAN, being a deep learning method, is capable of modelling complicated data distributions, enabling the effective modelling of a large variety of cars. Its major weakness is that as a global method, modelling the fine details in the local region is challenging. Lazy learning works well to preserve local features by generating a local subspace with relevant data samples. We demonstrate that the combined use of GAN and lazy learning produces is able to produce high-quality results, in which different types of cars with complicated local features can be generated effectively with a single sketch. Our method outperforms existing ones using other machine learning structures such as the variational autoencoder. Naoki Nozawa, Hubert P. H. Shum, Edmond S. L. Ho, Shigeo Morishima |
Vis. Comput. | 5 |
| 2021 | LineChaser: A Smartphone-Based Navigation System for Blind People to Stand in LinesabstractStanding in line is one of the most common social behaviors in public spaces but can be challenging for blind people. We propose an assistive system named LineChaser, which navigates a blind user to the end of a line and continuously reports the distance and direction to the last person in the line so that they can be followed. LineChaser uses the RGB camera in a smartphone to detect nearby pedestrians, and the built-in infrared depth sensor to estimate their position. Via pedestrian position estimations, LineChaser determines whether nearby pedestrians are standing in line, and uses audio and vibration signals to notify the user when they should start/stop moving forward. In this way, users can stay correctly positioned while maintaining social distance. We have conducted a usability study with 12 blind participants. LineChaser allowed blind participants to successfully navigate lines, significantly increasing their confidence in standing in lines. Masaki Kuribayashi, Seita Kayukawa, Hironobu Takagi, Chieko Asakawa, Shigeo Morishima |
CHI | 5 |
| 2021 | Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric LearningabstractThis paper describes a representation learning method for disentangling an arbitrary musical instrument sound into latent pitch and timbre representations. Although such pitch-timbre disentanglement has been achieved with a variational autoencoder (VAE), especially for a predefined set of musical instruments, the latent pitch and timbre representations are outspread, making them hard to interpret. To mitigate this problem, we introduce a metric learning technique into a VAE with latent pitch and timbre spaces so that similar (different) pitches or timbres are mapped close to (far from) each other. Specifically, our VAE is trained with additional contrastive losses so that the latent distances between two arbitrary sounds of the same pitch or timbre are minimized, and those of different pitches or timbres are maximized. This training is performed under weak supervision that uses only whether the pitches and timbres of two sounds are the same or not, instead of their actual values. This improves the generalization capability for unseen musical instruments. Experimental results show that the proposed method can find better-structured disentangled representations with pitch and timbre clusters even for unseen musical instruments. Keitaro Tanaka, Ryo Nishikimi, Yoshiaki Bando, Kazuyoshi Yoshii, Shigeo Morishima |
ICASSP | 5 |
| 2021 | Audio-Oriented Video Interpolation Using Key PoseabstractThis paper describes a deep learning-based method for long-term video interpolation that generates intermediate frames between two music performance videos of a person playing a specific instrument. Recent advances in deep learning techniques have successfully generated realistic images with high-fidelity and high-resolution in short-term video interpolation. However, there is still room for improvement in long-term video interpolation due to lack of resolution and temporal consistency of the generated video. Particularly in music performance videos, the music and human performance motion need to be synchronized. We solved these problems by using human poses and music features essential for music performance in long-term video interpolation. By closely matching human poses with music and videos, it is possible to generate intermediate frames that synchronize with the music. Specifically, we obtain the human poses of the last frame of the first video and the first frame of the second video in the performance videos to be interpolated as key poses. Then, our encoder–decoder network estimates the human poses in the intermediate frames from the obtained key poses, with the music features as the condition. In order to construct an end-to-end network, we utilize a differentiable network that transforms the estimated human poses in vector form into the human pose in image form, such as human stick figures. Finally, a video-to-video synthesis network uses the stick figures to generate intermediate frames between two music performance videos. We found that the generated performance videos were of higher quality than the baseline method through quantitative experiments. Takayuki Nakatsuka, Yukitaka Tsuchiya, Masatoshi Hamanaka, Shigeo Morishima |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2020 | Do We Need Sound for Sound Source Localization?
Takashi Oya, Shohei Iwase, Ryota Natsume, Takahiro Itazuri, Shugo Yamaguchi, Shigeo Morishima |
ACCV (6) | 6 |
| 2020 | Adversarial Knowledge Distillation for a Compact GeneratorabstractIn this paper, we propose memory-efficient Generative Adversarial Nets (GANs) in line with knowledge distillation. Most existing GANs have a shortcoming in terms of the number of model parameters and low processing speed. Here, to tackle the problem, we propose Adversarial Knowledge Distillation for Generative models (AKDG) for highly efficient GANs, in terms of unconditional generation. Using AKDG, model size and processing speed are substantively reduced. Through an adversarial training exercise with a distillation discriminator, a student generator successfully mimics a teacher generator in fewer model layers and fewer parameters and at a higher processing speed. Moreover, our AKDG is network architecture-agnostic. A Comparison of AKDG-applied models to vanilla models suggests that it achieves closer scores to a teacher generator and more efficient performance than a baseline method with respect to Inception Score (IS) and Frechet Inception Distance (FID). In CIFAR-10 experiments, improving IS/FID 1.17pt/55.19pt and in LSUN bedroom experiments, improving FID 71.1pt in comparison to the conventional distillation method for GANs. Our project page is https://maguro27.github.io/AKDG/. Hideki Tsunashima, Hirokatsu Kataoka, Junji Yamato, Qiu Chen, Shigeo Morishima |
ICPR | 5 |
| 2020 | Asynchronous Eulerian Liquid SimulationabstractAbstract We present a novel method for simulating liquid with asynchronous time steps on Eulerian grids. Previous approaches focus on Smoothed Particle Hydrodynamics (SPH), Material Point Method (MPM) or tetrahedral Finite Element Method (FEM) but the method for simulating liquid purely on Eulerian grids have not yet been investigated. We address several challenges specifically arising from the Eulerian asynchronous time integrator such as regional pressure solve, asynchronous advection, interpolation, regional volume preservation, and dedicated segregation of the simulation domain according to the liquid velocity. We demonstrate our method on top of staggered grids combined with the level set method and the semi‐Lagrangian scheme. We run several examples and show that our method considerably outperforms the global adaptive time step method with respect to the computational runtime on scenes where a large variance of velocity is present. Tatsuya Koike, Shigeo Morishima, Ryoichi Ando |
Comput. Graph. Forum | 2 |
| 2020 | Real-time rendering of layered materials with anisotropic normal distributionsabstractThis paper proposes a lightweight bidirectional scattering distribution function (BSDF) model for layered materials with anisotropic reflection and refraction properties. In our method, each layer of the materials can be described by a microfacet BSDF using an anisotropic normal distribution function (NDF). Furthermore, the NDFs of layers can be defined on tangent vector fields, which differ from layer to layer. Our method is based on a previous study in which isotropic BSDFs are approximated by projecting them onto base planes. However, the adequateness of this previous work has not been well investigated for anisotropic BSDFs. In this paper, we demonstrate that the projection is also applicable to anisotropic BSDFs and that the BSDFs are approximated by elliptical distributions using covariance matrices. Tomoya Yamaguchi 0002, Tatsuya Yatagawa, Yusuke Tokuyoshi, Shigeo Morishima |
Comput. Vis. Media | 4 |
| 2020 | Resolving hand-object occlusion for mixed reality with joint deep learning and model optimizationabstractAbstract By overlaying virtual imagery onto the real world, mixed reality facilitates diverse applications and has drawn increasing attention. Enhancing physical in‐hand objects with a virtual appearance is a key component for many applications that require users to interact with tools such as surgery simulations. However, due to complex hand articulations and severe hand‐object occlusions, resolving occlusions in hand‐object interactions is a challenging topic. Traditional tracking‐based approaches are limited by strong ambiguities from occlusions and changing shapes, while reconstruction‐based methods show a poor capability of handling dynamic scenes. In this article, we propose a novel real‐time optimization system to resolve hand‐object occlusions by spatially reconstructing the scene with estimated hand joints and masks. To acquire accurate results, we propose a joint learning process that shares information between two models and jointly estimates hand poses and semantic segmentation. To facilitate the joint learning system and improve its accuracy under occlusions, we propose an occlusion‐aware RGB‐D hand data set that mitigates the ambiguity through precise annotations and photorealistic appearance. Evaluations show more consistent overlays compared with literature, and a user study verifies a more realistic experience. Hubert P. H. Shum, Shigeo Morishima |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | Audio-visual object removal in 360-degree videosabstractAbstract We present a novel concept audio–visual object removal in 360-degree videos, in which a target object in a 360-degree video is removed in both the visual and auditory domains synchronously. Previous methods have solely focused on the visual aspect of object removal using video inpainting techniques, resulting in videos with unreasonable remaining sounds corresponding to the removed objects. We propose a solution which incorporates direction acquired during the video inpainting process into the audio removal process. More specifically, our method identifies the sound corresponding to the visually tracked target object and then synthesizes a three-dimensional sound field by subtracting the identified sound from the input 360-degree video. We conducted a user study showing that our multi-modal object removal supporting both visual and auditory domains could significantly improve the virtual reality experience, and our method could generate sufficiently synchronous, natural and satisfactory 360-degree videos. Ryo Shimamura, Yuki Koyama 0001, Takayuki Nakatsuka, Satoru Fukayama, Masahiro Hamasaki, Masataka Goto, Shigeo Morishima |
Vis. Comput. | 8 |
| 2020 | Data compression for measured heterogeneous subsurface scattering via scattering profile blendingabstractSubsurface scattering involves the complicated behavior of light beneath the surfaces of translucent objects that includes scattering and absorption inside the object’s volume. Physically accurate numerical representation of subsurface scattering requires a large number of parameters because of the complex nature of this phenomenon. The large amount of data restricts the use of the data on memory-limited devices such as video game consoles and mobile phones. To address this problem, this paper proposes an efficient data compression method for heterogeneous subsurface scattering. The key insight of this study is that heterogeneous materials often comprise a limited number of base materials, and the size of the subsurface scattering data can be significantly reduced by parameterizing only a few base materials. In the proposed compression method, we represent the scattering property of a base material using a function referred to as the base scattering profile. A small subset of the base materials is assigned to each surface position, and the local scattering property near the position is described using a linear combination of the base scattering profiles in the log scale. The proposed method reduces the data by a factor of approximately 30 compared to a state-of-the-art method, without significant loss of visual quality in the rendered graphics. In addition, the compressed data can also be used as bidirectional scattering surface reflectance distribution functions (BSSRDF) without incurring much computational overhead. These practical aspects of the proposed method also facilitate the use of higher-resolution BSSRDFs in devices with large memory capacity. Tatsuya Yatagawa, Hideki Todo, Yasushi Yamaguchi 0001, Shigeo Morishima |
Vis. Comput. | 4 |
| 2020 | LinSSS: linear decomposition of heterogeneous subsurface scattering for real-time screen-space rendering
Tatsuya Yatagawa, Yasushi Yamaguchi 0001, Shigeo Morishima |
Vis. Comput. | 3 |
| 2019 | BBeep: A Sonic Collision Avoidance System for Blind Travellers and Nearby PedestriansabstractWe present an assistive suitcase system, BBeep, for supporting blind people when walking through crowded environments. BBeep uses pre-emptive sound notifications to help clear a path by alerting both the user and nearby pedestrians about the potential risk of collision. BBeep triggers notifications by tracking pedestrians, predicting their future position in real-time, and provides sound notifications only when it anticipates a future collision. We investigate how different types and timings of sound affect nearby pedestrian behavior. In our experiments, we found that sound emission timing has a significant impact on nearby pedestrian trajectories when compared to different sound types. Based on these findings, we performed a real-world user study at an international airport, where blind participants navigated with the suitcase in crowded areas. We observed that the proposed system significantly reduces the number of imminent collisions. Seita Kayukawa, Keita Higuchi, João Guerreiro 0002, Shigeo Morishima, Yoichi Sato 0001, Kris Makoto Kitani, Chieko Asakawa |
CHI | 4 |
| 2019 | SiCloPe: Silhouette-Based Clothed PeopleabstractWe introduce a new silhouette-based representation for modeling clothed human bodies using deep generative models. Our method can reconstruct a complete and textured 3D model of a person wearing clothes from a single input picture. Inspired by the visual hull algorithm, our implicit representation uses 2D silhouettes and 3D joints of a body pose to describe the immense shape complexity and variations of clothed people. Given a segmented 2D silhouette of a person and its inferred 3D joints from the input picture, we first synthesize consistent silhouettes from novel view points around the subject. The synthesized silhouettes which are the most consistent with the input segmentation are fed into a deep visual hull algorithm for robust 3D shape prediction. We then infer the texture of the subject's back view using the frontal image and segmentation mask as input to a conditional generative adversarial network. Our experiments demonstrate that our silhouette-based model is an effective representation and the appearance of the back view can be predicted reliably using an image-to-image translation network. While classic methods based on parametric models often fail for single-view images of subjects with challenging clothing, our approach can still produce successful results, which are comparable to those obtained from multi-view input. Ryota Natsume, Shunsuke Saito, Zeng Huang, Weikai Chen 0001, Chongyang Ma, Hao Li 0015, Shigeo Morishima |
CVPR | 7 |
| 2019 | PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationabstractWe introduce Pixel-aligned Implicit Function (PIFu), an implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer both 3D surface and texture from a single image, and optionally, multiple input images. Highly intricate shapes, such as hairstyles, clothing, as well as their variations and deformations can be digitized in a unified way. Compared to existing representations used for 3D deep learning, PIFu produces high-resolution surfaces including largely unseen regions such as the back of a person. In particular, it is memory efficient unlike the voxel representation, can handle arbitrary topology, and the resulting surface is spatially aligned with the input image. Furthermore, while previous techniques are designed to process either a single image or multiple views, PIFu extends naturally to arbitrary number of views. We demonstrate high-resolution and robust reconstructions on real world images from the DeepFashion dataset, which contains a variety of challenging clothing types. Our method achieves state-of-the-art performance on a public benchmark and outperforms the prior work for clothed human digitization from a single image. Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Hao Li 0015, Angjoo Kanazawa |
ICCV | 4 |
| 2019 | Automatic Sign Dance Synthesis from Gesture-based Sign LanguageabstractAutomatic dance synthesis has become more and more popular due to the increasing demand in computer games and animations. Existing research generates dance motions without much consideration for the context of the music. In reality, professional dancers make choreography according to the lyrics and music features. In this research, we focus on a particular genre of dance known as sign dance, which combines gesture-based sign language with full body dance motion. We propose a system to automatically generate sign dance from a piece of music and its corresponding sign gesture. The core of the system is a Sign Dance Model trained by multiple regression analysis to represent the correlations between sign dance and sign gesture/music, as well as a set of objective functions to evaluate the quality of the sign dance. Our system can be applied to music visualization, allowing people with hearing difficulties to understand and enjoy music. Naoya Iwamoto, Hubert P. H. Shum, Wakana Asahina, Shigeo Morishima |
MIG | 4 |
| 2019 | 3D Car Shape Reconstruction from a Single Sketch ImageabstractEfficient car shape design is a challenging problem in both the automotive industry and the computer animation/games industry. In this paper, we present a system to reconstruct the 3D car shape from a single 2D sketch image. To learn the correlation between 2D sketches and 3D cars, we propose a Variational Autoencoder deep neural network that takes a 2D sketch and generates a set of multi-view depth & mask images, which are more effective representation comparing to 3D mesh, and can be combined to form the 3D car shape. To ensure the volume and diversity of the training data, we propose a feature-preserving car mesh augmentation pipeline for data augmentation. Since deep learning has limited capacity to reconstruct fine-detail features, we propose a lazy learning approach that constructs a small subspace based on a few relevant car samples in the database. Due to the small size of such a subspace, fine details can be represented effectively with a small number of parameters. With a low-cost optimization process, a high-quality car with detailed features is created. Experimental results show that the system performs consistently to create highly realistic cars of substantially different shape and topology, with a very low computational cost. Naoki Nozawa, Hubert P. H. Shum, Edmond S. L. Ho, Shigeo Morishima |
MIG | 4 |
| 2019 | Audio-Based Automatic Generation of a Piano Reduction Score by Considering the Musical Structure
Hirofumi Takamori, Takayuki Nakatsuka, Satoru Fukayama, Masataka Goto, Shigeo Morishima |
MMM (2) | 5 |
| 2019 | A Study on the Sense of Burden and Body Ownership on Virtual SlopeabstractThis paper provides insight into the burden when people are walking up and down slopes in a virtual environment (VE) while actually walking on a flat floor in the real environment (RE). In RE, we feel a physical load during walking uphill or downhill. To reproduce such physical load in the VE, we provided visual stimuli to users by changing their step length. In order to investigate how the stimuli affect a sense of burden and body ownership, we performed a user study where participants walked on slopes in the VE. We found that changing the step length has a significant impact on a burden on the user and less correlation between body ownership and step length. Ryo Shimamura, Seita Kayukawa, Takayuki Nakatsuka, Shoki Miyakawa, Shigeo Morishima |
VR | 5 |
| 2019 | Real-time Indirect Illumination of Emissive Inhomogeneous Volumes using Layered Polygonal Area LightsabstractAbstract Indirect illumination involving with visually rich participating media such as turbulent smoke and loud explosions contributes significantly to the appearances of other objects in a rendering scene. However, previous real‐time techniques have focused only on the appearances of the media directly visible from the viewer. Specifically, appearances that can be indirectly seen over reflective surfaces have not attracted much attention. In this paper, we present a real‐time rendering technique for such indirect views that involves the participating media. To achieve real‐time performance for computing indirect views, we leverage layered polygonal area lights (LPALs) that can be obtained by slicing the media into multiple flat layers. Using this representation, radiance entering each surface point from each slice of the volume is analytically evaluated to achieve instant calculation. The analytic solution can be derived for standard bidirectional reflectance distribution functions (BRDFs) based on the microfacet theory. Accordingly, our method is sufficiently robust to work on surfaces with arbitrary shapes and roughness values. In addition, we propose a quadrature method for more accurate rendering of scenes with dense volumes, and a transformation of the domain of volumes to simplify the calculation and implementation of the proposed method. By taking advantage of these computation techniques, the proposed method achieves real‐time rendering of indirect illumination for emissive volumes. Takahiro Kuge, Tatsuya Yatagawa, Shigeo Morishima |
Comput. Graph. Forum | 3 |
| 2018 | FSNet: An Identity-Aware Generative Model for Image-Based Face Swapping
Ryota Natsume, Tatsuya Yatagawa, Shigeo Morishima |
ACCV (6) | 3 |
| 2018 | Automatic paper summary generation from visual and textual informationabstractDue to the recent boom in artificial intelligence (AI) research, including computer vision (CV), it has become impossible for researchers in these fields to keep up with the exponentially increasing number of manuscripts. In response to this situation, this paper proposes the paper summary generation (PSG) task using a simple but effective method to automatically generate an academic paper summary from raw PDF data. We realized PSG by combination of vision-based supervised components detector and language-based unsupervised important sentence extractor, which is applicable for a trained format of manuscripts. We show the quantitative evaluation of ability of simple vision-based components extraction, and the qualitative evaluation that our system can extract both visual item and sentence that are helpful for understanding. After processing via our PSG, the 979 manuscripts accepted by the Conference on Computer Vision and Pattern Recognition (CVPR) 2018 are available1 . It is believed that the proposed method will provide a better way for researchers to stay caught with important academic papers. Shintaro Yamamoto, Yoshihiro Fukuhara, Ryota Suzuki 0006, Shigeo Morishima, Hirokatsu Kataoka |
ICMV | 4 |
| 2018 | Face Retrieval Framework Relying on User's Visual MemoryabstractThis paper presents an interactive face retrieval framework for clarifying an image representation envisioned by a user. Our system is designed for a situation in which the user wishes to find a person but has only visual memory of the person. We address a critical challenge of image retrieval across the user's inputs. Instead of target-specific information, the user can select several images (or a single image) that are similar to an impression of the target person the user wishes to search for. Based on the user's selection, our proposed system automatically updates a deep convolutional neural network. By interactively repeating these process (human-in-the-loop optimization), the system can reduce the gap between human-based similarities and computer-based similarities and estimate the target image representation. We ran user studies with 10 subjects on a public database and confirmed that the proposed framework is effective for clarifying the image representation envisioned by the user easily and quickly. Yugo Sato, Tsukasa Fukusato, Shigeo Morishima |
ICMR | 3 |
| 2018 | Resolving occlusion for 3D object manipulation with hands in mixed realityabstractDue to the need to interact with virtual objects, the hand-object interaction has become an important element in mixed reality (MR) applications. In this paper, we propose a novel approach to handle the occlusion of augmented 3D object manipulation with hands by exploiting the nature of hand poses combined with tracking-based and model-based methods, to achieve a complete mixed reality experience without necessities of heavy computations, complex manual segmentation processes or wearing special gloves. The experimental results show a frame rate faster than real-time and a great accuracy of rendered virtual appearances, and a user study verifies a more immersive experience compared to past approaches. We believe that the proposed method can improve a wide range of mixed reality applications that involve hand-object interactions. Hubert P. H. Shum, Shigeo Morishima |
VRST | 3 |
| 2018 | Thickness-aware voxelizationabstractAbstract Voxelization is a crucial process for many computer graphics applications such as collision detection, rendering of translucent objects, and global illumination. However, in some situations, although the mesh looks good, the voxelization result may be undesirable. In this paper, we describe a novel voxelization method that uses the graphics processing unit for surface voxelization. Our improvements on the voxelization algorithm can address a problem of state‐of‐the‐art voxelization, which cannot deal with thin parts of the mesh object. We improve the quality of voxelization on both normal mediation and surface correction. Furthermore, we investigate our voxelization methods on indirect illumination, showing the improvement on the quality of real‐time rendering. Zhuopeng Zhang, Shigeo Morishima, Changbo Wang |
Comput. Animat. Virtual Worlds | 2 |
| 2018 | High-fidelity facial reflectance and geometry inference from an unconstrained imageabstractWe present a deep learning-based technique to infer high-quality facial reflectance and geometry given a single unconstrained image of the subject, which may contain partial occlusions and arbitrary illumination conditions. The reconstructed high-resolution textures, which are generated in only a few seconds, include high-resolution skin surface reflectance maps, representing both the diffuse and specular albedo, and medium- and high-frequency displacement maps, thereby allowing us to render compelling digital avatars under novel lighting conditions. To extract this data, we train our deep neural networks with a high-quality skin reflectance and geometry database created with a state-of-the-art multi-view photometric stereo system using polarized gradient illumination. Given the raw facial texture map extracted from the input image, our neural networks synthesize complete reflectance and displacement maps, as well as complete missing regions caused by occlusions. The completed textures exhibit consistent quality throughout the face due to our network architecture, which propagates texture features from the visible region, resulting in high-fidelity details that are consistent with those seen in visible regions. We describe how this highly underconstrained problem is made tractable by dividing the full inference into smaller tasks, which are addressed by dedicated neural networks. We demonstrate the effectiveness of our network design with robust texture completion from images of faces that are largely occluded. With the inferred reflectance and geometry data, we demonstrate the rendering of high-fidelity 3D avatars from a variety of subjects captured under different lighting conditions. In addition, we perform evaluations demonstrating that our method can infer plausible facial reflectance and geometric details comparable to those obtained from high-end capture devices, and outperform alternative approaches that require only a single unconstrained input image. Shugo Yamaguchi, Shunsuke Saito, Koki Nagano, Weikai Chen 0001, Kyle Olszewski, Shigeo Morishima, Hao Li 0015 |
ACM Trans. Graph. | 7 |
| 2017 | Voice Animator: Automatic Lip-Synching in Limited Animation by Audio
Shoichi Furukawa, Tsukasa Fukusato, Shugo Yamaguchi, Shigeo Morishima |
ACE | 4 |
| 2017 | DanceDJ: A 3D Dance Animation Authoring System for Live Performance
Naoya Iwamoto, Takuya Kato, Hubert P. H. Shum, Ryo Kakitsuka, Kenta Hara 0001, Shigeo Morishima |
ACE | 6 |
| 2017 | Facial video age progression considering expression changeabstractThis paper proposes an age progression method for facial videos. Age is one of the main factors that changes the appearance of the face, due to the associated sagging, spots, and wrinkles. These aging features change in appearance depending on facial expressions. As an example, we see wrinkles appear in the face of the young when smiling, but the shape of wrinkles changes in older faces. Previous work has not considered the temporal changes of the face, using only static images as databases. To solve this problem, we extend the texture synthesis approach to use facial videos as source videos. First, we spatio-temporally align the videos of database to match the sequence of a target video. Then, we synthesize an aging face and apply the temporal changes of the target age to the wrinkles appearing in the facial expression image in the target video. As a result, our method successfully generates expression changes specific to the target age. Shintaro Yamamoto, Pavel A. Savkin, Takuya Kato, Shoichi Furukawa, Shigeo Morishima |
CGI | 5 |
| 2017 | Outside-in monocular IR camera based HMD pose estimation via geometric optimizationabstractAccurately tracking a Head Mounted Display (HMD) with a 6 degree of freedom is essential to achieve a comfortable and a nausea free experience in Virtual Reality. Existing commercial HMD systems using synchronized Infrared (IR) camera and blinking IR-LEDs can achieve highly accurate tracking. However, most of the off-the-shelf cameras do not support frame synchronization. In this paper, we propose a novel method for real time HMD pose estimation without using any camera synchronization or LED blinking. We extended over the state of the art pose estimation algorithm by introducing geometrically constrained optimization. In addition, we propose a novel system to increase robustness to the blurred IR-LEDs patterns appearing at high-velocity movements. The quantitative evaluations showed significant improvements in pose stability and accuracy over wide rotational movements as well as a decrease in runtime. Pavel A. Savkin, Shunsuke Saito, Jarich Vansteenberge, Tsukasa Fukusato, Lochlainn Wilson, Shigeo Morishima |
VRST | 6 |
| 2016 | Perception of drowsiness based on correlation with facial image featuresabstractThis paper presents a video-based method for detecting drowsiness. Generally, human beings can perceive their fatigue and drowsiness through looking at faces. The ability to perceive the fatigue and the drowsiness has been studied in many ways. The drowsiness detection method based on facial videos has been proposed [Nakamura et al. 2014]. In their method, a set of the facial features calculated with the Computer Vision techniques and the k-nearest neighbor algorithm are applied to classify drowsiness degree. However, the facial features that are ineffective against reproducing the perception of human beings with the machine learning method are not removed. This factor can decrease the detection accuracy. Yugo Sato, Takuya Kato, Naoki Nozawa, Shigeo Morishima |
SAP | 4 |
| 2016 | A soundtrack generation system to synchronize the climax of a video clip with musicabstractIn this paper, we present a soundtrack generation system that can automatically add a soundtrack with the length and climax points aligned to those of a video clip. Adding a soundtrack to a video clip is an important process in video editing. Editors tend to add chorus sections to the climax points of the video clip by replacing and concatenating musical segments. However, this process is time-consuming. Our system automatically detects climaxes of both the video clips and music based on feature extraction and analysis. This enables the system to add a soundtrack in which the climax is synchronized to the climax of the video clip. We evaluated the generated soundtracks through a subjective evaluation. Haruki Sato, Tatsunori Hirai, Tomoyasu Nakano, Masataka Goto, Shigeo Morishima |
ICME | 5 |
| 2016 | Computational Cartoonist: A Comic-Style Video Summarization System for Anime Films
Tsukasa Fukusato, Tatsunori Hirai, Shunya Kawamura, Shigeo Morishima |
MMM (1) | 4 |
| 2016 | MusicMixer: Automatic DJ System Considering Beat and Latent Topic Similarity
Tatsunori Hirai, Hironori Doi, Shigeo Morishima |
MMM (1) | 3 |
| 2016 | Frame-Wise Continuity-Based Video Summarization and Stretching
Tatsunori Hirai, Shigeo Morishima |
MMM (1) | 2 |
| 2015 | MusicMixer: computer-aided DJ system based on an automatic song mixingabstractIn this paper, we present MusicMixer, a computer-aided DJ system that helps DJs, specifically with song mixing. MusicMixer continuously mixes and plays songs using an automatic music mixing method that employs audio similarity calculations. By calculating similarities between song sections that can be naturally mixed, MusicMixer enables seamless song transitions. Though song mixing is the most fundamental and important factor in DJ performance, it is difficult for untrained people to seamlessly connect songs. MusicMixer realizes automatic song mixing using an audio signal processing approach; therefore, users can perform DJ mixing simply by selecting a song from a list of songs suggested by the system, enabling effective DJ song mixing and lowering entry barriers for the inexperienced. We also propose personalization for song suggestions using a preference memorization function of MusicMixer. Tatsunori Hirai, Hironori Doi, Shigeo Morishima |
Advances in Computer Entertainment | 3 |
| 2015 | FOCUSING PATCH: Automatic Photorealistic Deblurring for Facial Images by Patch-Based Color Transfer
Masahide Kawai, Shigeo Morishima |
MMM (1) | 2 |
| 2015 | Facial Aging Simulator by Data-Driven Component-Based Texture Cloning
Daiki Kuwahara, Akinobu Maejima, Shigeo Morishima |
MMM (2) | 3 |
| 2015 | Affective Music Recommendation System Based on the Mood of Input Video
Shoto Sasaki, Tatsunori Hirai, Hayato Ohya, Shigeo Morishima |
MMM (2) | 4 |
| 2015 | Multi-layer Lattice Model for Real-Time Dynamic Character DeformationabstractDue to the recent advancement of computer graphics hardware and software algorithms, deformable characters have become more and more popular in real-time applications such as computer games. While there are mature techniques to generate primary deformation from skeletal movement, simulating realistic and stable secondary deformation such as jiggling of fats remains challenging. On one hand, traditional volumetric approaches such as the finite element method require higher computational cost and are infeasible for limited hardware such as game consoles. On the other hand, while shape matching based simulations can produce plausible deformation in real-time, they suffer from a stiffness problem in which particles either show unrealistic deformation due to high gains, or cannot catch up with the body movement. In this paper, we propose a unified multi-layer lattice model to simulate the primary and secondary deformation of skeleton-driven characters. The core idea is to voxelize the input character mesh into multiple anatomical layers including the bone, muscle, fat and skin. Primary deformation is applied on the bone voxels with lattice-based skinning. The movement of these voxels is propagated to other voxel layers using lattice shape matching simulation, creating a natural secondary deformation. Our multi-layer lattice framework can produce simulation quality comparable to those from other volumetric approaches with a significantly smaller computational cost. It is best to be applied in real-time applications such as console games or interactive animation creation. Naoya Iwamoto, Hubert P. H. Shum, Longzhi Yang, Shigeo Morishima |
Comput. Graph. Forum | 4 |
| 2014 | VRMixer: mixing video and real world with video segmentationabstractThis paper presents VRMixer, a system that mixes real world and a video clip letting a user enter the video clip and realize a virtual co-starring role with people appearing in the clip. Our system constructs a simple virtual space by allocating video frames and the people appearing in the clip within the user's 3D space. By measuring the user's 3D depth in real time, the time space of the video clip and the user's 3D space become mixed. VRMixer automatically extracts human images from a video clip by using a video segmentation technique based on 3D graph cut segmentation that employs face detection to detach the human area from the background. A virtual 3D space (i.e., 2.5D space) is constructed by positioning the background in the back and the people in the front. In the video clip, the user can stand in front of or behind the people by using a depth camera. Real objects that are closer than the distance of the clip's background will become part of the constructed virtual 3D space. This synthesis creates a new image in which the user appears to be a part of the video clip, or in which people in the clip appear to enter the real world. We aim to realize "video reality," i.e., a mixture of reality and video clips using VRMixer. Tatsunori Hirai, Satoshi Nakamura 0002, Tsubasa Yumura, Shigeo Morishima |
Advances in Computer Entertainment | 4 |
| 2014 | Automatic depiction of onomatopoeia in animation considering physical phenomenaabstractThis paper presents a method that enables the estimation and depiction of onomatopoeia in computer-generated animation based on physical parameters. Onomatopoeia is used to enhance physical characteristics and movement, and enables users to understand animation more intuitively. We experiment with onomatopoeia depiction in scenes within the animation process. To quantify onomatopoeia, we employ Komatsu's [2012] assumption, i.e., onomatopoeia can be expressed by n-dimensional vector. We also propose phonetic symbol vectors based on the correspondence of phonetic symbols to the impressions of onomatopoeia using a questionnaire-based investigation. Furthermore, we verify the positioning of onomatopoeia in animated scenes. The algorithms directly combine phonetic symbols to estimate optimum onomatopoeia. They use a view-dependent Gaussian function to display onomatopoeias in animated scenes. Our method successfully recommends optimum onomatopoeias using only physical parameters, so that even amateur animators can easily create onomatopoeia animation. Tsukasa Fukusato, Shigeo Morishima |
MIG | 2 |
| 2014 | Macroscopic and microscopic deformation coupling in up-sampled cloth simulationabstractABSTRACT Various methods of predicting the deformation of fine‐scale cloth from coarser resolutions have been explored. However, the influence of fine‐scale deformation has not been considered in coarse‐scale simulations. Thus, the simulation of highly nonhomogeneous detailed cloth is prone to large errors. We introduce an effective method to simulate cloth made of nonhomogeneous, anisotropic materials. We precompute a macroscopic stiffness that incorporates anisotropy from the microscopic structure, using the deformation computed for each unit strain. At every time step of the simulation, we compute the deformation of coarse meshes using the coarsened stiffness, which saves computational time and add higher‐level details constructed by the characteristic displacement of simulated meshes. We demonstrate that anisotropic and inhomogeneous cloth models can be simulated efficiently using our method. © 2014 The Authors. Computer Animation and Virtual Worlds published by John Wiley & Sons, Ltd. Shunsuke Saito, Nobuyuki Umetani, Shigeo Morishima |
Comput. Animat. Virtual Worlds | 3 |
| 2013 | Real-time Hair Simulation on Mobile DeviceabstractHair rendering and simulation is a fundamental part in the representation of virtual characters. But intensive calculation for the dynamic on thousands of hair strands makes the task much challengeable, especially on a portable device. The aim of this short paper is to solve the problem of how to perform real-time hair simulation and rendering on mobile device. In this paper, the process of hair simulation and rendering is adapted according to the property of mobile device hardware. To increase the number of hair strands of simulation, we adopted the Dynamic follow-the-leader (DFTL) method and altered it by our new method of interpolation. We also pictured a rendering strategy basing on the survey of the limitation of mobile GPU. Lastly we present an innovational method that carried out order independent transparency at a relatively inexpensive cost. Zhuopeng Zhang, Shigeo Morishima |
MIG | 2 |
| 2012 | Fast-accurate 3D face model generation using a single video camera
Tomoya Hara, Hiroyuki Kubo, Akinobu Maejima, Shigeo Morishima |
ICPR | 4 |
| 2010 | Automatic generation of head models and facial animations considering personal characteristicsabstractWe propose a new automatic head modeling system to generate individualized head models which can express person-specific facial expressions. The head modeling system consists of two core processes. The head modeling process with the proposed automatic mesh completion generates a whole head model only from facial range scan data. The key shape generation process generates key shapes for the generated head model based on physics-based facial muscle simulation with an individual muscle layout estimated from subject's facial expression videos. Facial animations considering personal characteristics can be synthesized using the individualized head model and key shapes. Experimental results show that the proposed system can generate head models where 84% of subjects can identify themselves. Therefore, we conclude that our head modeling system is effective to games and entertainment systems like a Future Cast System. Akinobu Maejima, Hiroto Yarimizu, Hiroyuki Kubo, Shigeo Morishima |
VRST | 4 |
| 2010 | The effects of virtual characters on audiences' movie experienceabstractIn this paper, we first present a new audience-participating movie form in which 3D virtual characters of audiences are constructed by computer graphics (CG) technologies and are embedded into a in a pre-rendered movie as different roles. Then, we investigate how the audiences respond to these virtual characters using physiological and subjective evaluation methods. To facilitate the investigation, we present three versions of a movie to an audience—a Traditional version, its SDIM version with the participation of the audience’s virtual character, and its SFDIM version with the co-participation of the audience and her/his friends’ virtual characters. The subjective evaluation results show that the participation of virtual characters indeed causes increased subjective sense of spatial presence and engagement, and emotional reaction; moreover, SFDIM performs significantly better than SDIM, due to the co-participation of friends’ virtual characters. Also, we find that the audiences experience not only significantly different galvanic skin response (GSR) changes on average—changing trend over time and number of fluctuations—but they also show the increased phasic GSR responses to the appearance of their own or friends’ virtual 3D characters on the screen. The evaluation results demonstrate the success of the new audience-participating movie form and contribute to understanding how people respond to virtual characters in a role-playing entertainment interface. Tao Lin 0006, Shigeo Morishima, Akinobu Maejima, Ningjiu Tang |
Interact. Comput. | 2 |
| 2009 | Automatic voice assignment tool for Instant Casting movie SystemabstractIn Instant Casting movie System, a personal CG character is automatically generated. The character resembles a participant in a face geometry and texture. However, the voice of a character was an alternative voice determined by the gender of the participant. Therefore sometimes it's not enough to identify the personality of a CG character. In this paper, an automatic pre-scored voice assignment tool for a personal CG character is presented. Voice is essential to identify a personal character as well as a face feature. Our proposed system selects the most similar voice to the participants from voice database, and assigns it as a voice of CG character. Voice similarity criterion is presented by combination of eight acoustic features. After assigning voice data to a personal character, the voice track is played back in synchronization with the movement of the CG character. 60 voice variations have been prepared to our voice database. Validity of the assigned voice has been evaluated by MOS value. The proposed method has achieved 68% of the theoretical figure that is calculated by preliminary experiments. Yoshihiro Adachi, Shinichi Kawamoto, Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2009 | Dive into the movie: an instant casting and immersive experience in the storyabstractOur research project, Dive into Movie (DIM) aims to build a new genre of interactive entertainment which enables anyone to easily participate in a movie by assuming a role and enjoying an embodied, first-hand theater experience. This is specifically accomplished by replacing the original roles of the precreated traditional movie with user created, high-realism, 3-D CG characters. DIM movie is in some sense a hybrid entertainment form, somewhere between a game and storytelling. We hope that DIM movies might enhance interaction and offer more dramatic presence, engagement, and fun for the audience. In DIM movie, audiences can experience highrealism 3-D CG character action with individualized facial characteristics, expression, gait and voice. The DIM system has two key features: First, it can full-automatically create a CG character in a few minutes from capturing the face, body, gait and voice feature of a user and generating her/his corresponding CG animation, to inserting the individualized CG character into the movie in real-time which do not cause any discomfort to the participant; Second, the DIM system makes it possible for multiple participants to take part in a movie at the same time in different roles, such as a family, a circle of friends, etc. And also DIM project proposes a panoramic image capture/projection system and 3D sound field capture/playback system from performer, Äôs view and standing point. Shigeo Morishima |
VRST | 1 |
| 2009 | Interactive shadowing for 2D AnimeabstractAbstract In this paper, we propose an instant shadow generation technique for 2D animation, especially Japanese Anime. In traditional 2D Anime production, the entire animation including shadows is drawn by hand so that it takes long time to complete. Shadows play an important role in the creation of symbolic visual effects. However shadows are not always drawn due to time constraints and lack of animators especially when the production schedule is tight. To solve this problem, we develop an easy shadowing approach that enables animators to easily create a layer of shadow and its animation based on the character's shapes. Our approach is both instant and intuitive. The only inputs required are character or object shapes in input animation sequence with alpha value generally used in the Anime production pipeline. First, shadows are automatically rendered on a virtual plane by using a Shadow Map1based on these inputs. Then the rendered shadows can be edited by simple operations and simplified by the Gaussian Filter. Several special effects such as blurring can be applied to the rendered shadow at the same time. Compared to existing approaches, ours is more efficient and effective to handle automatic shadowing in real‐time. Copyright © 2009 John Wiley & Sons, Ltd. Eiji Sugisaki, Seah Hock Soon, Feng Tian 0006, Shigeo Morishima |
Comput. Animat. Virtual Worlds | 4 |
| 2008 | Using subjective and physiological measures to evaluate audience-participating movie experienceabstractIn this paper we subjectively and physiologically investigate the effects of the audiences' 3D virtual actor in a movie on their movie experience, using the audience-participating movie DIM as the object of study. In DIM, the photo-realistic 3D virtual actors of audience are constructed by combining current computer graphics (CG) technologies and can act different roles in a pre-rendered CG movie. To facilitate the investigation, we presented three versions of a CG movie to an audience---a Traditional version, its Self-DIM (SDIM) version with the participation of the audience's virtual actor, and its Self-Friend-DIM (SFDIM) version with the co-participation of the audience and his friends' virtual actors. The results show that the participation of audience's 3D virtual actors indeed cause increased subjective sense of presence and engagement, and emotional reaction; moreover, SFDIM performs significantly better than SDIM, due to increased social presence. Interestingly, when watching the three movie versions, subjects experienced not only significantly different galvanic skin response (GSR) changes on average---changing trend over time, and number of fluctuations---but they also experienced phasic GSR increase when watching their own and friends' virtual 3D actors appearing on the movie screen. These results suggest that the participation of the 3D virtual actors in a movie can improve interaction and communication between audience and the movie. Tao Lin 0006, Akinobu Maejima, Shigeo Morishima |
AVI | 3 |
| 2008 | Perceptual similarity measurement of speech by combination of acoustic featuresabstractFuture cast system is a new entertainment system where participant’s face is captured and rendered into the movie as an instant Computer Graphics (CG) movie star, which had been first exhibited at the 2005 World Exposition in Aichi Japan. We are working to add new functionality which enables mapping not only faces but also speech individualities to the cast. Our approach is to find a speaker with the closest speech individuality and apply voice conversion. This paper investigates acoustic features to estimate perceptual similarity of speech individuality. We propose a method linearly combined eight acoustic features related to the perception of speech individualities. The proposed method optimizes weights for the acoustic features considering perceptual similarities. We have evaluated performance of our method with Spearman’s rank correlation coefficients to perceptual similarities. As the results, the experiments evidenced that the proposed method achieves a correlation coefficient of 0.66. Yoshihiro Adachi, Shinichi Kawamoto, Shigeo Morishima, Shun Nakamura |
ICASSP | 3 |
| 2008 | Post-recording tool for instant casting movie systemabstractThis paper proposes a universal user-friendly post-recording tool for an Instant Casting Movie System (ICS) that enables anyone to be a movie star using his or her own voice and faces. A personal CG character is automatically generated by scanning one's face geometry and image in ICS. Voice is as essential to identify a person as face. However, a character's voice is only based on gender in ICS. We proposed a novel voice recording tool for participants of all ages in a short time. Post-recording tasks are very difficult because speakers should speak in synchronization with the mouth movements of the CG characters. Therefore this task is generally recorded by professional voice actors. Our proposed tool has the following four features: 1) various supporting information for synchronization with voice and mouth movement timing for users; 2) automatic post-processing of recorded voices for compositing mixed audio; 3) intuitively displays operation for people of all ages; and 4) handles multiple users in parallel for quick recording. We developed a prototype speech synchronization system using a post-recording tool and conducted subjective evaluation experiments of it. Over 60% of the subjects responded that the tool's interface can be operated easily. Shinichi Kawamoto, Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ACM Multimedia | 3 |
| 2006 | Hair motion cloning from cartoon animation sequencesabstractAbstract This paper describes a new approach to create cartoon hair animation that allows users to use existing cel character animation sequences. We demonstrate the generation of cartoon hair animation accentuated in ‘anime‐like’ motions. The novelty of this method is that users can choose the existing cel animation of a character's hair animation and apply environmental elements such as wind to other characters with a three‐dimensional structure. In fact, users can reuse existing cartoon sequences as input to endow another character with environmental elements as if both characters exist in the same scene. A three‐dimensional character's hair motions are created based on hair motions from input cartoon animation sequences. First, users extract hair shapes at each frame from input sequences from which they then construct physical equations. ‘Anime‐like’ hair motion is created by using these physical equations. Copyright © 2006 John Wiley & Sons, Ltd. Eiji Sugisaki, Yosuke Kazama, Shigeo Morishima, Natsuko Tanaka, Akiko Sato |
Comput. Animat. Virtual Worlds | 3 |
| 2003 | Model-based talking face synthesis for anthropomorphic spoken dialog agent systemabstractTowards natural human-machine communication, interface technologies by way of speech and image information have been intensively developed. An anthropomorphic dialog agent is an ideal system, which integrates spoken dialog and natural facial expressions. This paper reports on our project aiming to create a general-purpose toolkit for building an easily customizable anthropomorphic agent. There have been almost no tools so far such as intuitive, easy to understand, fully interactive, and open source. Our anthropomorphic agent is designed to fulfill these requirements. This toolkit consists four modules, multi modal dialog integration, speech recognition, speech synthesis, and face image synthesis. These modules are highly modularized and interlinked by a simple communication protocols.In this paper, we focus on the construction of an agent's face image synthesis. For this part lip movement control synchronous to the speech signal and facial emotion expression are the most important parts. We developed the face image synthesis module (FSM) that only requires one frontal face image, and can be used by any skill level of users. A user's original agent can be generated by easy adjustment of the frontal face image and the generic wire-frame model. The paper describes overall system diagram and specifically the agent's face image synthesis part. Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ACM Multimedia | 2 |
| 2003 | How to capture absolute human skeletal postureabstractCommercially available motion capture products give us fairly precise movements of human body segments but do not measure enough information to define skeletal posture in its entirety. This sketch describes how to obtain the complete posture of skeletal structure with the help of marker locations relative to bones that are derived from MRI data sets. Shoichiro Iwasawa, Kiyoshi Kojima, Kenji Mase, Shigeo Morishima |
SIGGRAPH | 4 |
| 2002 | Audio-visual speech translation with automatic lip syncqronization and face tracking based on 3-D head modelabstractSpeech-to-speech translation has been studied to realize natural human communication beyond language barriers. Toward further multi-modal natural communication, visual information such as face and lip movements will be necessary. In this paper, we introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. We conduct subjective evaluation by connected digit discrimination using data with and without audiovisual lip-synchronicity. The results confirm the sufficient quality of the proposed audio-visual translation system. Shigeo Morishima, Shin Ogata, Kazumasa Murai, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2002 | Multi-Modal Translation System and Its EvaluationabstractSpeech-to-speech translation has been studied to realize natural human communication beyond language barriers. Toward further multi-modal natural communication, visual information such as face and lip movements will be necessary. We introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. We conduct subjective evaluation tests using the connected digit discrimination test using data with and without audio-visual lip-synchronization. The results confirm the significant quality of the proposed audio-visual translation system and the importance of lip-synchronization. Shigeo Morishima, Satoshi Nakamura 0001 |
ICMI | 1 |
| 2002 | HyperMask - projecting a talking head onto a real object
Tatsuo Yotsukura, Shigeo Morishima, Frank Nielsen, Kim Binsted, Claudio S. Pinhanez |
Vis. Comput. | 2 |
| 2001 | Hypermask: talking head projected onto moving surfaceabstractHypermask is a system which projects an animated face onto a physical mask, worn by an actor. As the mask moves within a prescribed area, its position and orientation are detected by a camera, and the projected image changes with respect to the viewpoint of the audience. The lips of the projected face are automatically synthesized in real time with the voice of the actor, who also controls the facial expressions. As a theatrical tool, Hypermask enables a new style of storytelling. As a prototype system, we propose to put a self-contained Hypermask system in a trolley (disguised as a linen cart), so that it projects onto the mask worn by the actor pushing the trolley. Tatsuo Yotsukura, Shigeo Morishima |
ICIP (3) | 2 |
| 2001 | Automatic Face Tracking And Model Match-Move In Video Sequence Using 3d Face Model
Takafumi Misawa, Kazumasa Murai, Satoshi Nakamura 0001, Shigeo Morishima |
ICME | 4 |
| 2001 | Trends of Learning Technology Standard
Shigeo Morishima, Shin Ogata, Satoshi Nakamura 0001 |
ICME | 1 |
| 2001 | Model-Based Lip Synchronization With Automatically Translated Systhetic Voice Toward A Multi-Modal Translation SystemabstractIn this paper, we introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. Shin Ogata, Kazumasa Murai, Satoshi Nakamura 0001, Shigeo Morishima |
ICME | 4 |
| 2000 | Human Body Postures from Trinocular Camera ImagesabstractThis paper proposes a new real-time method for estimating human postures in 3D from trinocular images. In this method, an upper body orientation detection and a heuristic contour analysis are performed on the human silhouettes extracted from the trinocular images so that representative points such as the top of the head can be located. The major joint positions are estimated based on a genetic algorithm-based learning procedure. 3D coordinates of the representative points and joints are then obtained from the two views by evaluating the appropriateness of the three views. The proposed method implemented on a personal computer runs in real-time. Experimental results show high estimation accuracies and the effectiveness of the view selection process. Shoichiro Iwasawa, Jun Ohya, Kazuhiko Takahashi, Tatsumi Sakaguchi, Shigeo Morishima, Kazuyuki Ebihara |
FG | 5 |
| 1999 | Multi-Media Ambiance Communication Based on Actual Moving PicturesabstractMulti-media ambiance communication refers to a means of shared-space communication, that makes use of actual moving pictures captured by video camera, combined with the laws of perspective, as used in painting and the visual characteristics of human beings, to establish a photo-realistic quality three-dimensional image space that users can naturally feel part of. We aim to enable this ambiance communication by basing the shared-space on actual moving pictures rather than on an accurate three-dimensional image space, such as is used in computer graphics. Tadashi Ichikawa, Tetsuya Yoshimura, Kunio Yamada, Toshifumi Kanamaru, Hiromichi Suga, Shoichiro Iwasawa, Takeshi Naemura, Kiyoharu Aizawa, Shigeo Morishima, Takahiro Saito |
ICIP (3) | 9 |
| 1999 | Face-To-Face Communicative Avatar Driven by VoiceabstractRecently computer can make cyberspace to walk through by an interactive virtual reality technique. An avatar in cyberspace can bring us a virtual face-to-face communication environment. In this paper we realize an avatar which has a real face in cyberspace to construct a multi-user communication system by voice transmission through network. Voice from microphone is transmitted and analyzed, then mouth shape and facial expression of avatar are synchronously estimated and synthesized on real time. And also we introduce an entertainment application of a real-time voice driven synthetic face. This project is named "Fifteen Seconds of Fame" which is an example of interactive movie. Shigeo Morishima, Tatsuo Yotsukura |
ICIP (3) | 1 |
| 1998 | 3D Estimation of Facial Muscle Parameter from the 2D Marker Movement Using Neural Network
Takahiro Ishikawa, Hajime Sera, Shigeo Morishima, Demetri Terzopoulos |
ACCV (2) | 3 |
| 1998 | Facial Image Reconstruction by Estimated Muscle Parameter
Takahiro Ishikawa, Hajime Sera, Shigeo Morishima, Demetri Terzopoulos |
FG | 3 |
| 1998 | Real-Time Human Posture Estimation Using Monocular Thermal Images
Shoichiro Iwasawa, Kazuyuki Ebihara, Jun Ohya, Shigeo Morishima |
FG | 4 |
| 1998 | Facial muscle parameter decision from 2D frontal imageabstractMuscle based face image synthesis is one of the most realistic approaches to realizing life-like agents in a computer. A facial muscle model is composed of facial tissue elements and muscles. In this model, forces are calculated effecting facial tissue elements by contraction of each muscle strength, so the combination of each muscle parameter decides a specific facial expression. Each muscle parameter is decided based on a trial and error procedure comparing the sample photograph and generated image using our Muscle-Editor to generate a specific face image. We propose a strategy of automatic estimation of facial muscle parameters from 2D marker movements using a neural network. We can also carry out 3D motion estimation from 2D point or flow information in a captured image under restriction of a physics based face model. Shigeo Morishima, Takahiro Ishikawa, Demetri Terzopoulos |
ICPR | 1 |
| 1997 | Real-Time Estimation of Human Body Posture from Monocular Thermal ImagesabstractThis paper introduces a new real-time method to estimate the posture of a human from thermal images acquired by an infrared camera regardless of the back-ground and lighting conditions. Distance transformation is performed for the human body area extracted from the thresholded thermal image for the. Calculation of the center of gravity. After the orientation of the upper half of the body is obtained by calculating the moment of inertia, significant points such as the top of the head, the tips of the hands and foot are heuristically located. In addition, the elbow and foot positions are estimated from the detected (significant) points using a genetic algorithm based learning procedure. The experimental results demonstrate the robustness of the proposed algorithm and real-time (faster than 20 frames per second) performance. Shoichiro Iwasawa, Kazuyuki Ebihara, Jun Ohya, Shigeo Morishima |
CVPR | 4 |
| 1996 | Face feature extraction from spatial frequency for dynamic expression recognitionabstractA new facial feature extraction technique for expression recognition is proposed. We employ the spatial frequency domain information to obtain robust performance to the random noise on a image or the lighting conditions. It exhibited high ability sufficiently even if combined with a low-performance region tracking method. As an application of this technique, we have constructed a dynamic facial expression recognition system. We use hidden Markov models to utilize temporal changes in the facial expressions. The spatial frequency information and the temporal information make better rates of facial expression recognition. In the experiment, we established a correct response rate of approximately 84.1% of recognition with six categories. Tatsumi Sakaguchi, Shigeo Morishima |
ICPR | 2 |
| 1991 | Speech-to-image media conversion based on VQ and neural networkabstractAutomatic media conversion schemes from speech to a facial image and a construction of a real-time image synthesis system are presented. The purpose of this research is to realize an intelligent human-machine interface or intelligent communication system with synthesized human face images. A human face image is reconstructed on the display of a terminal using a 3-D surface model and texture mapping technique. Facial motion images are synthesized by transformation of the 3-D model. In the motion driving method, based on vector quantization and the neural network, the synthesized head image can appear to speak some given words and phrases naturally, in synchronization with voice signals from a speaker.> Shigeo Morishima, Hiroshi Harashima |
ICASSP | 1 |
| 1991 | A Media Conversion from Speech to Facial Image for Intelligent Man-Machine InterfaceabstractAn automatic field motion image synthesis scheme (driven by speech) and a real-time image synthesis design are presented. The purpose of this research is to realize an intelligent human-machine interface or intelligent communication system with talking head images. A human face is reconstructed on the display of a terminal using a 3-D surface model and texture mapping technique. Facial motion images are synthesized naturally by transformation of the lattice points on 3-D wire frames. Two driving motion methods, a text-to-image conversion scheme and a voice-to-image conversion scheme, are proposed. In the first method, the synthesized head image can appear to speak some given words and phrases naturally. In the second case, some mouth and jaw motions can be synthesized in synchronization with voice signals from a speaker. Facial expressions other than mouth shape and jaw position can be added at any moment, so it is easy to make the facial model appear angry, to smile, to appear sad, etc., by special modification rules. These schemes were implemented on a parallel image computer system. A real-time image synthesizer was able to generate facial motion images on the display at a TV image video rate.> Shigeo Morishima, Hiroshi Harashima |
IEEE J. Sel. Areas Commun. | 1 |
| 1990 | Real-time facial action image synthesis system driven by speech and textabstractAutomatic facial motion image synthesis schemes and a real-time system design are presented. The purpose of this schemes is to realize an intelligent human-machine interface or intelligent communication system with talking head images. Human's face is reconstructed with 3D surface model and texture mapping technique on the display of terminal. Facial motion images are synthesized naturally by transformation of the lattice points on wire frames. Two types of motion drive methods, text to image conversion and speech to image conversion are proposed in this paper. In the former manner, synthesized head can speak some given texts naturally and in the latter case, some mouth and jaw motions can be synthesized in time to speech signal of behind speaker. These schemes were implemented to a parallel image computer and a real-time image synthesizer could output facial motion images to the display as fast as video rate. Shigeo Morishima, Kiyoharu Aizawa, Hiroshi Harashima |
VCIP | 1 |
| 1989 | An intelligent facial image coding driven by speech and phonemeabstractThe authors propose and compare two types of model-based facial motion coding schemes, i.e. synthesis by rules and synthesis by parameters. In synthesis by rules, facial motion images are synthesized on the basis of rules extracted by analysis of training image samples that include all of the phonemes and coarticulation. This system can be utilized as an automatic facial animation synthesizer from text input or as a man-machine interface using the facial motion image. In synthesis by parameters, facial motion images are synthesized on the basis of a code word index of speech parameters. Experimental results indicate good performance for both systems, which can create natural facial-motion images with very low transmission rate. Details of 3-D modeling, algorithm synthesis, and performance are discussed.> Shigeo Morishima, Kiyoharu Aizawa, Hiroshi Harashima |
ICASSP | 1 |
| 1986 | A proposal of a knowledge based isolated word recognitionabstractThis paper describes a knowledge based isolated Japanese word recognition algorithm. The program is written with Prolog/KR [4] and has two basic inference processes, i.e., a Bottom-up search and a Top-down search. In the Bottom-up process, a segmentation and a vowel decision are performed and some target word patterns are generated. The Top-down process includes a consonant decision using a score of each candidate word calculated based on the Fuzzy Set Theory [1]. In the vowel inference, a template matching is applied mainly. In the segmentation, heuristic rules based on the spectrum transition and the wave form are used. But in the consonant inference, each rule has a hierarchy structure and it is defined automatically in the form of the multi-valued threshold function from learning data. This system can treat an obscure information about the consonant classification and to select the most effective decision rule in order to simplify the understanding process. The truth rate of consonant recognition is better than using statistical method. Shigeo Morishima, Hiroshi Harashima, Hiroshi Miyakawa |
ICASSP | 1 |