Takao Nakamura

dblp:03/6228 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
11since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance Estimation
abstract
This study investigated differences in the contributions of various features to in-person (IP) and video-mediated (VM) meetings. We focused on estimating important utterances using both an IP and a VM meeting corpora as the analysis data. A transformer model with dialogue history was used to estimate important utterances, and five types of input (text, speaker’s audio, others’ audio, speaker’s video, and others’ video) were fed to the model. A comparison of the models for IP and VM revealed that the speaker’s audio has a strong effect on the IP model, the video of the other participants strongly affects the VM model, and the text and others’ audio strongly affects both models in estimating important utterances.
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Atsushi Fukayama, Takao Nakamura
ICASSP5
2023 Prediction of Love-Like Scores After Speed Dating Based on Pre-obtainable Personal Characteristic Information
Ryo Ishii, Fumio Nihei, Yoko Ishii, Atsushi Otsuka, Kazuya Matsuo, Narichika Nomoto, Atsushi Fukayama, Takao Nakamura
INTERACT (4)8
2023 How Far ahead Can Model Predict Gesture Pose from Speech and Spoken Text?
abstract
We investigated how far into the future nonverbal behavior can be predicted from speech and speech text. Specifically, we build a model that generates future behaviors from speech and speech text information and evaluate the quality of the generated behaviors. This helps to clarify how far into the future behavior can be accurately predicted. Our experimental results show that in Gesture Pose Generation using speech and speech text, on the basis of the input speech and text, the nonverbal behavior up to at least 500 ms ahead can be predicted with objective evaluation values that are the same as those when no future prediction is made. This result shows a new possibility for Gesture Pose Generation using speech and speech text to predict the future up to at least 500 ms ahead with no performance degradation.
Ryo Ishii, Akira Morikawa, Shin'ichiro Eitoku, Atsushi Fukayama, Takao Nakamura
IVA5
2023 A Study of Prediction of Listener's Comprehension Based on Multimodal Information
abstract
During dialogues, speakers need to be able to predict whether their partners understand their message. This is important for not only for human-to-human interaction but also human-to-agent interaction. We consider that if the listener's comprehension level can be automatically predicted, interactive agents will be able to communicate appropriately according to the user's comprehension level. However, to the best of our knowledge, there is no case study that reveals how comprehension can be predicted based on multimodal information about the listener. In this study, we attempt to predict comprehension levels on the basis of the listener's multimodal information. First, we construct a dialogue corpus consisting of the listener's comprehension levels and the listener's multimodal information. Next, we construct machine learning models that predict the listener's comprehension levels on the basis of the listener's multimodal information. Our results suggest that our model was able to predict a listener's comprehension level on the basis of a listener's multimodal information. In addition, two movements, the lifting of the cheeks and the pulling up of the corners of the lips, were suggested to be important in assessing the listener's level of comprehension.
Shunichi Kinoshita, Toshiki Onishi, Naoki Azuma, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA6
2023 Prediction of Various Backchannel Utterances Based on Multimodal Information
abstract
The listener's backchannels are an important part of dialogues. With appropriate backchannels, people are able to smoothly promote dialogues. Thus, backchannels are considered to be important in dialogues between not only humans but also humans and agents. Progress has been made in studying dialogue agents that perform natural affable dialogue. However, we have not clarified whether the listener's various backchannel types are predictable using the speaker's multimodal information. In this paper, we attempt to predict a listener's various backchannel types on the basis of the speaker's multimodal information in dialogues. First, we construct a dialogue corpus that consists of multimodal information of a speaker's utterances and a listener's backchannels. Second, we construct machine learning models to predict a listener's various backchannel types on the basis of a speaker's multimodal information. Our results suggest that our model was able to predict a listener's various backchannel types on the basis of a speaker's multimodal information.
Toshiki Onishi, Naoki Azuma, Shunichi Kinoshita, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
IVA6
2022 Effect of repetitive motion intervention on self-avatar on the sense of self-individuality
abstract
In recent years, the human Digital Twin has been discussed as new technology. When we discuss a world in which one’s self-avatar autonomously performs social activities in cyberspace, the questions arise whether or not the behavior of the avatars feels like one’s own, and whether or not we can approve of the self-avatars’ social activities on behalf of ourselves. We define such feeling as the sense of self-individuality. In this study, we focused on the situation in which self-avatars perform presentations on behalf of ourselves to investigate the effect of the modification experience on the presentation motions by self-avatars on the sense of self-individuality. We conducted VR-based experiments in which the motion modification intervention was performed on self-avatars over eight weeks by 24 experiment participants. As a result, we found that the sense of self-individuality was improved as the number of modifications and interventions increased. However, we found that the intensity of motion modification did not correlate with the improvement of the sense of self-individuality in this experiment condition. We also found that the sense of self-individuality was reduced when others intervened in the motion. From these results, we clarified that the experience of motion modification on self-avatars is significant when designing the behavior of avatars acting on behalf of ourselves in human Digital Twin. Further investigation is required to clarify the effect of the long-term intervention on behavior to distinguish between the mere exposure effect.
Tetsunari Inamura, Shin'ichiro Eitoku, Iwaki Toshima, Shinya Shimizu, Atsushi Fukayama, Shiro Ozawa, Takao Nakamura
HAI7
2022 Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura
INTERSPEECH7
2022 Analysis of praising skills focusing on utterance contents
Asahi Ogushi, Toshiki Onishi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
INTERSPEECH6
2022 Determining most suitable listener backchannel type for speaker's utterance
abstract
A major hurdle in achieving a dialogue system that enables smooth dialogue is to determine how to generate an appropriate response to a user's utterance. Previous research has focused mainly on estimating whether to make an utterance backchannel in response to the user's utterance. We go one step further by examining, for the first time, the relationship between the type of utterance backchannel to be used and intent and type of the speaker's utterance, known as a dialogue act (DA). Specifically, we propose a new method for classifying utterance backchannels into nine types. We also created a corpus consisting of the DAs of speaker utterances and the backchannel types of listener utterances then used it to analyze the relationship between a speaker's and listener's utterances. Our findings clarify that the occurrence frequencies of a listener's backchannel types significantly depend on the DAs of the speaker's utterances. Since the goal of our research is to construct a dialogue system that generates a more natural backchannel, this classification method, which determines certain types of aids from the speaker's DA, will be beneficial to such a system.
Akira Morikawa, Ryo Ishii, Hajime Noto, Atsushi Fukayama, Takao Nakamura
IVA5
2022 A Comparison of Praising Skills in Face-to-Face and Remote Dialogues
abstract
Praising behavior is considered to an important method of communication in daily life and social activities. An engineering analysis of praising behavior is therefore valuable. However, a dialogue corpus for this analysis has not yet been developed. Therefore, we develop corpuses for face-to-face and remote two-party dialogues with ratings of praising skills. The corpuses enable us to clarify how to use verbal and nonverbal behaviors for successfully praise. In this paper, we analyze the differences between the face-to-face and remote corpuses, in particular the expressions in adjudged praising scenes in both corpuses, and also evaluated praising skills. We also compare differences in head motion, gaze behavior, facial expression in high-rated praising scenes in both corpuses. The results showed that the distribution of praising scores was similar in face-to-face and remote dialogues, although the ratio of the number of praising scenes to the number of utterances was different. In addition, we confirmed differences in praising behavior in face-to-face and remote dialogues.
Toshiki Onishi, Asahi Ogushi, Yohei Tahara, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, Akihiro Miyata
LREC6
2021 How People Distinguish Individuals from their Movements: Toward the Realization of Personalized Agents
abstract
Demands for agents that replicate the characteristics of specific individuals are increasing. Although ways to implement personality traits into the virtual agents’ movement have been widely researched, ways to create aspects of individuality that can be identified as belonging to specific individuals have not. To clarify how well humans can identify individuals from short movements and what elements of movement contribute to the perception of individuality, we examined the relationship between the degree of confidence in personal identification and statics of gesture movement. In the experiment, participants were asked to compare pairs of short presentation animations and give their degree of confidence that the two animations were of the same person. The animations were created with motion data from performers and with 3D-CG characters to reduce the differences in appearances and shot angles. We calculated five expressivity parameters from the wrist movement for each gesture and compared the answers from the participants. The results showed that the participants were able to distinguish individuals doing the same action and recognize the individuals by the spatial and temporal extents of their movement, which were represented by how much space they use and how fast they moved their wrists. This study clarifies the cognitive aspects of what elements need to be reproduced to develop agents with individuality.
Chihiro Takayama, Mitsuhiro Goto, Shin'ichiro Eitoku, Ryo Ishii, Hajime Noto, Shiro Ozawa, Takao Nakamura
HAI7
2013 GaN Substrate Technologies for Optical Devices
abstract
Large GaN single-crystal substrates with low dislocation density are the key materials for the commercial production of GaN-based laser diodes. We developed a new method to reduce the dislocations, named dislocation elimination by the epitaxial-growth with inverse-pyramidal pits (DEEP). A thick GaN film is epitaxially grown on a GaAs substrate with hydride vapor-phase epitaxy and then is separated from the GaAs substrate. The thick GaN layer grows with numerous large inverse-pyramidal pits. As the growth proceeds, dislocations in the GaN film are concentrated to the center of the pit and a wide area with low dislocation density is formed within the pit except the center area. To control the dislocation artificially, the position of the pits is fixed at a predetermined position by means of the selective growth of different polarity GaN film. This process was named as advanced DEEP (A-DEEP). GaN substrates with the A-DEEP method satisfied all the requirements for the violet laser diodes. We continued to develop new GaN substrates such as large-diameter c-plane substrates as well as nonpolar/semipolar GaN substrates. In particular, we overcame the green gap problem and developed the world's first true green laser diodes by selecting an optimal crystal plane.
Takao Nakamura, Kensaku Motoki
Proc. IEEE1
2007 A Fast, Robust Watermark Detection Scheme for Videos Captured on Camera Phones
abstract
Digital watermarking technology can be applied to reference services that provide digital information related to pictures taken by mobile phone cameras. Conventional reference services for mobile phones are designed for still images and detect watermarks from single pictures. We propose a video watermarking scheme for fast, robust detection to develop a reference service for video. To handle spatial synchronization, the scheme uses the rectangular tracking technique, and to provide temporal synchronization and robustness against noise, it uses the watermark technique based on the SFPSS method. We also introduce a quantitative evaluation method to determine reliability of the detection results. Finally, we evaluate the scheme's implementation on a mobile phone and demonstrate its efficiency.
Takao Nakamura, Susumu Yamamoto, Ryo Kitahara, Atsushi Katayama, Takayuki Yasuno, Noboru Sonehara
ICME1
2006 Fast Watermark Detection Scheme from Camera-captured Images on Mobile Phones
abstract
Digital watermarking technology would be very useful as part of a related service introduction system (RSIS); which links physical objects in the real world, such as printed photographs, to those in the cyber world. In this paper, we focus on a camera-equipped mobile phone as an RSIS terminal, and propose a fast watermark detection scheme for the captured images. The proposed scheme consists of two processes; to correct geometric distortion of the captured image, and to detect watermark information from the corrected image. For our scheme, we also propose a fast quadrangle detection algorithm and a robust watermarking algorithm. Moreover, we introduce a quantitative evaluation method for determining the reliability of watermark detection, which is essential for RSIS services. Finally, we show that the proposed scheme enables users to detect embedded information far less than one second, even when implemented as a Java application on a mobile phone with limited resources, and our experiments confirm the efficiency of our scheme.
Takao Nakamura, Atsushi Katayama, Masashi Yamamuro, Noboru Sonehara
Int. J. Pattern Recognit. Artif. Intell.1
2004 New high-speed frame detection method: Side Trace Algorithm (STA) for i-appli on cellular phones to detect watermarks
abstract
We developed a system that enables a camera-equipped cellular phone to read digital watermarks embedded in various media in real time, and that presents to the user a link to a Web page, video, or music associated with that watermark information. A picture captured by a camera is the result of applying a projective transformation combining rotation, scaling, and tilting to the original picture, The picture must be subjected to an inverse projective transformation prior to reading the watermark in order to return it to the same geometric form as the original picture. This inverse transformation requires transformation parameters, and the corners of the picture outline can be used as feature points for determining these parameters. In this paper, we propose a Side Trace Algorithm (STA) that reduces the processing time required to find corners of the picture less than 1/100 that when using the Hough transform and the conventional pattern matching, and present results of its implementation.
Atsushi Katayama, Takao Nakamura, Masashi Yamamuro, Noboru Sonehara
MUM2
2004 Fast watermark detection scheme for camera-equipped cellular phone
abstract
Digital watermarking technology would be very useful as part of a related service introduction system (RSIS); this system provides related information to content, and the function of watermark in RSIS is analogous to that of barcode, i.e., watermark binds content ID to analog content such as an image on printed material.In this paper, we focus on a camera-equipped cellular phone used as a terminal for RSIS, and propose a fast watermark detection scheme from a captured image. The proposed scheme consists of two processes, one is to correct geometric distortion of the captured image, and the other is to detect watermark information from the rectified image. We also propose a new watermarking algorithm which is robust against small geometric distortion and suitable for the proposed scheme. Moreover, we introduce a quantitative evaluation method for indicating detection reliability, which is indispensable for RSIS service.Finally, we show that the proposed scheme enables users to detect embedded information in approximately one second, even when implemented as a Java application on a cell phone with limited resources, and report experiments that confirm the proposed scheme's efficiency.
Takao Nakamura, Atsushi Katayama, Masashi Yamamuro, Noboru Sonehara
MUM1
2002 Additional content-related service/product offering system based on new standards: MPEG-21 and content ID/DOI
abstract
This paper discusses the configuration of an Internet-based system in which products and services relating to digital content are offered to users by employing the ISO MPEG-21 standard and the de facto standards: Content ID (defined by cIDf) and DOI (defined by IDF).
Hideki Sakamoto, Masanori Yamada, Takao Nakamura, Tadashi Nakanishi
ICME (2)3
1986 Architecture of high-speed 22-bit floating-point digital signal processor
abstract
A high-speed floating-point DSP(MSM6992) has been successfully developed as a second-generation DSP. The DSP employs a 22-bit(16E6) floating-point data format and realizes a 100-ns instruction cycle as well as an arithmetic capability of 20- MFLOPS. Furthermore, the DSP features an excellent expandability and allows the expansion of both the program and data memories up to a maximum of 64K words. In addition, the DSP enables flexible system constructions, including multi-DSP systems. The DSP is implemented by a 2µm CMOS process and contains 125K transistors within a single chip. The dimensions of the chip are 12.7 × 10.6mm2and its power dissipation is 400mW. The chip is packaged in a 132-pin PGA. This paper describes the architecture of the high-speed floating-point DSP and its VLSI implementation.
Yoshikazu Mori, Toshio Jufuku, Masao Iida, Akira Nomura, Noboru Ichiura, Takao Nakamura
ICASSP6