Hiroki Tanaka

dblp:66/2476 · DBLP profile ↗
← Back
29ranked-venue papers
14as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 11 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Adaptive virtual agent: Design and evaluation for real-time human-agent interaction
abstract
When we converse, we adapt our behaviors to our interlocutors. The adaptation can serve to indicate our engagement which can also elicit enhancement of the involvement of others. Virtual agents (or socially interactive virtual agents) that play the role of interaction partners can improve the human users’ interaction experience by displaying continuous and adaptive behaviors in real time. Virtual agents have been used in multiple domains to improve user interaction and performance. The promising results of the endowment of adaptation to agents in increasing the agents’ perception and user experience were shown in previous studies. In this paper, we develop an adaptive virtual agent that renders real-time adaptive behaviors based on the behaviors shown by its human interlocutor. The ASAP model rendering reciprocally adaptive agent behavior was employed to realize the system. The system consists of four main parts: perception of social signals, agent adaptive behavior generation, agent visualization (i.e. rendering of the agent’s verbal and nonverbal behavior), and communication of signals. To showcase the usefulness of our adaptive agent, as a proof-of-concept we choose the e-health application of cognitive behavior therapy (CBT), which identifies and rectifies biased and irrational thoughts (or automatic thoughts). Through this study, we show the importance of giving the agent reciprocal adaptation capability notably in enhancing the user experience and the effectiveness of the CBT session. We validate the importance of endowing such adaptation capability by studying the difference between agents that are reciprocal adaptive, solely expressive (with mismatched behavior), and inexpressive (in a still posture) via questionnaires and measures related to the agent perception (naturalness, human-likeliness, synchrony, and engagement) for user experience and the CBT effectiveness (mood, anxiety, stress, and cognitive change). These results highlight the value of making virtual agents adapt in real time. This could lead to agents being capable of providing more personalized and interactive experiences for a wide range of applications. Also, we have collected a new human-agent interaction (HAI) database, HAI-CBT database, which is publicly available to the research community.
Jieyeon Woo, Kazuhiro Shidara, Catherine Achard, Hiroki Tanaka, Satoshi Nakamura 0001, Catherine Pelachaud
Int. J. Hum. Comput. Stud.4
2023 Acceptability and Trustworthiness of Virtual Agents by Effects of Theory of Mind and Social Skills Training
abstract
We constructed a social skills training system using virtual agents and developed a new training module for four basic tasks: declining, requesting, praising, and listening. Previous work demonstrated that a virtual agent's theory of mind influences the building of trust between agents and users. The purpose of this study is to explore the effect of trustworthiness, acceptability, familiarity, and likeability on the agents' theory of mind and the social skills training contents. In our experiment, 29 participants rated the trustworthiness and acceptability of the virtual agent after watching a video that featured levels of theory of mind and social skills training. Their ratings were obtained using self-evaluation measures at each stage. We confirmed that our users' trust and acceptability of the virtual agent were significantly changed depending on the level of the virtual agent's theory of mind. We also confirmed that the users' trust and acceptability in the trainer tended to improve after the social skills training.
Hiroki Tanaka, Takeshi Saga, Kota Iwauchi, Satoshi Nakamura 0001
FG1
2023 Multimodal Voice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conversation
abstract
Predicting the timing of utterances in dyadic conversations is essential for achieving natural interactions between humans and virtual agents. Since the former often use non-verbal cues to adjust the order of their speech, this study proposes a multimodal model incorporating non-verbal features using a Transformer-based voice activity prediction model. First, in line with previous research, we reproduced a baseline model that utilized audio features (audio waveform, voice activity frame, and voice activity history) as inputs. To this baseline model, we added non-verbal features: gaze direction, action units, head pose, and articular points. We compared our multimodal model with the baseline model to investigate the impact of non-verbal cues on voice activity prediction. We utilized a dyadic expert-novice conversation dataset and evaluated the average outcomes across ten model trainings. Results revealed that our proposed models with all the features improved the accuracy of the next speaker prediction by 2.3% and back-channeled prediction by 1.8% (p-value < 0.025). In particular, action units may contribute significantly to the turn-shift and back-channeled predictions. This study demonstrates that including non-verbal features in Transformer-based turn-taking models enhances the efficacy of models for predicting voice activity in dyadic conversations.
Kazuyo Onishi, Hiroki Tanaka, Satoshi Nakamura 0001
HAI2
2022 Propagation Characteristics of 920 MHz band LPWA for Inspections of Power Transmission Lines Using UAVs
abstract
This study considers an autonomous inspection system of power transmission lines using multicopter-type un-manned aerial vehicles (UAVs). In this system, base stations (BSs) are placed on power transmission towers, and the control information for autonomous flight of UAVs is exchanged between the BSs and UAVs through wireless transmissions. In this study, we conduct transmission experiments between the transmitter of a UAV and a receiver placed on the power transmission tower. In the transmission experiments, we employ LoRa in 920 MHz band and measure the received signal strength indicator (RSSI) to investigate propagation characteristics. From the measured results, we calculate the mean and standard deviation of RSSI and discuss the feasibility of the system.
Tekkan Okuda, Hiraku Okada, Hiroki Tanaka, Toshihiro Fujiwara, Nobuo Matsui, Chedlia Ben Naila, Masaaki Katayama
CCNC3
2022 3rd Workshop on Social Affective Multimodal Interaction for Health (SAMIH)
abstract
This workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user’s behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces.
Hiroki Tanaka, Satoshi Nakamura 0001, Kazuhiro Shidara, Jean-Claude Martin, Catherine Pelachaud
ICMI1
2021 Clustering of Human Movement Trajectories based on Distributional Representations Derived from Bi-directional LSTM Network with Geographical Coordinates
abstract
As the ubiquity of such wearable devices as smart-phones continues to deepen its presence in modern societies, it has become possible to analyze and visualize people who are moving as part of a trajectory of big data. In this study, we cluster human movement trajectories using time-series distributional representations. For the clustering, we calculated the distance of the representation vectors derived from neural network models. Previous work leveraged the Long short-term memory (LSTM) network to train the next mesh prediction. In this study, we propose using the Bi-directional LSTM (Bi-LSTM) network and the integrated additional geographical coordinates (latitude and longitude information) in models to accurately predict the next mesh and construct user clusters. As a result, we improved the accuracy of the next mesh prediction and obtained and visualized clusters of human movement trajectories.
Hiroki Tanaka, Takeshi Saga, Satoshi Nakamura 0001
IEEE BigData1
2021 2nd Workshop on Social Affective Multimodal Interaction for Health (SAMIH)
abstract
This workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user’s behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces.
Hiroki Tanaka, Satoshi Nakamura 0001, Jean-Claude Martin, Catherine Pelachaud
ICMI1
2021 Band-Stop Bandwidths Adjustment for a Periodic Disturbance Observer
abstract
A periodic disturbance deteriorates precision of motion control of machines. Suppression of the periodic disturbance is necessary so that the machines can move precisely. This paper proposes a periodic disturbance observer with adjustable band-stop bandwidths, which realizes wide band-stop bandwidths suppressing a frequency-varying periodic disturbance. Mean gain of the sensitivity function at harmonic frequencies is defined as a cost function, and the cost function designs an almost optimal proposed method by determining a parameter that adjust the band-widths. We validated the almost optimal proposed method that performed the best suppression performance under six variations of a fundamental frequency of a periodic disturbance through experiments.
Hiroki Tanaka, Hisayoshi Muramatsu
IECON1
2021 Experimental Evaluation of a Wireless LAN Relay System Using Unmanned Aerial Vehicles
abstract
A wireless LAN relay system is designed to extend the transmission range while maintaining stable communication. In this paper, we propose a wireless LAN relay system using multicopter unmanned aerial vehicles (UAVs) and evaluate the communication performance. For the demonstration experiment, we employ air-to-air and ground-to-air communications, where the former is a link between UAVs, and the latter is a link between a UAV and ground node. We measure the throughput and the received signal strength indicator, and clarify the communication performance of the wireless LAN relay system.
Hiroki Tanaka, Nobuo Matsui, Yuta Watanabe, Hiraku Okada, Masaaki Katayama
VTC Spring1
2020 Social Affective Multimodal Interaction for Health
abstract
This workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user's behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces.
Hiroki Tanaka, Satoshi Nakamura 0001, Jean-Claude Martin, Catherine Pelachaud
ICMI1
2020 Combining Audio and Brain Activity for Predicting Speech Quality
Ivan Halim Parmonangan, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH2
2019 Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor Decomposition
abstract
Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG research with speech production, a method using three-mode tensor decomposition (time x space x frequency) is proposed to perform speech artifact removal. Tensor decomposition enables simultaneous inspection of multiple modes, which suits the multi-way nature of EEG data. In a picture-naming task, we collected raw data with speech artifacts by placing two electrodes near the mouth to record lip EMG. Based on our evaluation, which calculated the correlation values between grand-averaged speech artifacts and the lip EMG, tensor decomposition outperformed the former methods that were based on independent component analysis (ICA) and blind source separation (BSS), both in detecting speech artifact (0.985) and producing clean data (0.101). Our proposed method correctly preserved the components unrelated to speech, which was validated by computing the correlation value between the grand-averaged raw data without EOG and cleaned data before the speech onset (0.92-0.94).
Holy Lovenia, Hiroki Tanaka, Sakriani Sakti, Ayu Purwarianti, Satoshi Nakamura 0001
ICASSP2
2019 Speech Quality Evaluation of Synthesized Japanese Speech Using EEG
Ivan Halim Parmonangan, Hiroki Tanaka, Sakriani Sakti, Shinnosuke Takamichi, Satoshi Nakamura 0001
INTERSPEECH2
2018 Graph Regularized Tensor Factorization for Single-Trial EEG Analysis
abstract
This study proposes a tensor factorization algorithm for electroencephalographies (EEGs) that incorporates the geometric structure of the electrode location. The purpose is removing noise caused by EEG activities which are irrelevant to stimuli presented to a subject from single-trial event-related potential (ERP) data. Canonical polyadic decomposition (CPD) is extended by adding a regularization term that controls the spatial smoothness of the decomposed components on a scalp. An initialization method using geometrical information is also proposed. The geometric structure of an EEG signal is expressed as an undirected graph where the similarities between electrodes are defined by their relative distances on a scalp. The effectiveness is demonstrated in a noise-removing experiment using pseudo-ERP, where the proposed method achieved better performance than the conventional CPD.
Hayato Maki, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001
ICASSP2
2018 Listening Skills Assessment through Computer Agents
abstract
Social skills training, performed by human trainers, is a well-established method for obtaining appropriate skills in social interaction. Previous work automated the process of social skills training by developing a dialogue system that teaches social skills through interaction with a computer agent. Even though previous work that simulated social skills training considered speaking skills, human social skills trainers take into account other skills such as listening. In this paper, we propose assessment of user listening skills during conversation with computer agents toward automated social skills training. We recorded data of 27 Japanese graduate students interacting with a female agent. The agent spoke to the participants about a recent memorable story and how to make a telephone call, and the participants listened. Two expert external raters assessed the participants' listening skills. We manually extracted features relating to eye fixation and behavioral cues of the participants, and confirmed that a simple linear regression with selected features can correctly predict a user's listening skills with above 0.45 correlation coefficient.
Hiroki Tanaka, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001
ICMI1
2018 Detection of Dementia from Responses to Atypical Questions Asked by Embodied Conversational Agents
Tsuyoki Ujiro, Hiroki Tanaka, Hiroyoshi Adachi, Hiroaki Kazui, Manabu Ikeda, Takashi Kudo, Satoshi Nakamura 0001
INTERSPEECH2
2018 Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags
Koichiro Yoshino, Hiroki Tanaka, Kyoshiro Sugiyama, Makoto Kondo, Satoshi Nakamura 0001
LREC2
2017 Tracking liking state in brain activity while watching multiple movies
abstract
Emotion is a valuable information in various applications ranging from human-computer interaction to automated multimedia content delivery. Conventional methods to recognize emotion were based on speech prosody cues, facial expression, and body language. However, this information may not appear when people watch a movie. In recent years, some studies have started to use electroencephalogram (EEG) signals in recognizing emotion. But, the EEG data were entirely analyzed in each scene of movies for emotion classification. Thus, the detailed information of emotional state changes cannot be extracted. In this study, we utilize EEG to track affective state during watching multiple movies. Experiments were done by measuring continuous liking state during watching three types of movies, and then constructing subject dependent emotional state tracking model. We used support vector machine (SVM) as a classifier, and support vector regression (SVR) for regression. As a result, the best classification accuracy was 77.6%, and the best regression model achieved 0.645 of correlation coefficient between actual liking state and predicted liking state. These results demonstrate that continuous emotional state can be predicted by our EEG-based method.
Naoto Terasawa, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001
ICMI2
2017 Subject-Independent Classification of Japanese Spoken Sentences by Multiple Frequency Bands Phase Pattern of EEG Response During Speech Perception
Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001
INTERSPEECH2
2017 Virtual Co-Eating: Making Solitary Eating Experience More Enjoyable
Monami Takahashi, Hiroki Tanaka, Hayato Yamana, Tatsuo Nakajima
ICEC2
2016 Personalized unknown word detection in non-native language reading using eye gaze
abstract
This paper proposes a method to detect unknown words during natural reading of non-native language text by using eye-tracking features. A previous approach utilizes gaze duration and word rarity features to perform this detection. However, while this system can be used by trained users, its performance is not sufficient during natural reading by untrained users. In this paper, we 1) apply support vector machines (SVM) with novel eye movement features that were not considered in the previous work and 2) examine the effect of personalization. The experimental results demonstrate that learning using SVMs and proposed eye movement features improves detection performance as measured by F-measure and that personalization further improves results.
Rui Hiraoka, Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001
ICMI2
2016 Automatic detection of very early stage of dementia through multimodal interaction with computer avatars
abstract
This paper proposes a new approach to detecting very early stage of dementia automatically. We develop a computer avatar with spoken dialog functionalities that produces natural spoken queries referring to Mini Mental State Examination, Wechsler Memory Scale-Revised and other related questions. Multimodal interactive data of spoken dialogues from 18 participants (9 dementias and 9 healthy controls) are recorded, and audiovisual features are extracted. We confirm that the support vector machines can classify into two groups with 0.94 detection performance as measured by areas under ROC curve. It is found that our system has possibilities to detect very early stage of dementia through spoken dialog with our computer avatars.
Hiroki Tanaka, Hiroyoshi Adachi, Norimichi Ukita, Takashi Kudo, Satoshi Nakamura 0001
ICMI1
2016 Teaching Social Communication Skills Through Human-Agent Interaction
abstract
There are a large number of computer-based systems that aim to train and improve social skills. However, most of these do not resemble the training regimens used by human instructors. In this article, we propose a computer-based training system that follows the procedure of social skills training (SST), a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. We attempt to automate the process of SST by developing a dialogue system named the automated social skills trainer , which teaches social communication skills through human-agent interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users. Its design is based on conventional SST performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. We performed a series of three experiments investigating (1) the advantages of using computer-based training systems compared to human-human interaction (HHI) by subjectively evaluating nervousness, ease of talking, and ability to talk well; (2) the relationship between speech language features and human social skills; and (3) the effect of computer-based training using our proposed system. Results of our first experiment show that interaction with an avatar decreases nervousness and increases the user's subjective impression of his or her ability to talk well compared to interaction with an unfamiliar person. The experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with social skills. Finally, experiments measuring the effect of performing SST with the proposed application show that participants significantly improve their skill, as assessed by separate evaluators, by using the system for 50 minutes. A user survey also shows that the users thought our system is useful and easy to use, and that interaction with the avatar felt similar to HHI.
Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001
ACM Trans. Interact. Intell. Syst.1
2015 Automated Social Skills Trainer
abstract
Social skills training is a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. In this paper, we attempt to automate the process of social skills training by developing a dialogue system named "automated social skills trainer," which provides social skills training through human-computer interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users to improve their social skills. Its design is based on conventional social skills training performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. An experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with autistic traits. Additional experiments measuring the effect of performing social skills training with the proposed application show that most participants improve their skill by using the system for 50 minutes.
Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001
IUI1
2014 Classification of social laughter in natural conversational speech
Hiroki Tanaka, Nick Campbell 0001
Comput. Speech Lang.1
2012 Fault Recovery Technique for TMR Softcore Processor System Using Partial Reconfiguration
Makoto Fujino, Hiroki Tanaka, Yoshihiro Ichinomiya, Motoki Amagasaki, Morihiro Kuga, Masahiro Iida, Toshinori Sueyoshi
ICA3PP (1)2
2008 Performance Evaluation of Cooperative Relaying Networks Using 3D Ray Launching Method for Wireless Propagation Prediction
abstract
Cooperative relaying is a promising technique for multihop wireless networks to exploit spatial diversity. Most of the studies of multihop relaying assumed a simple i.i.d. Rayleigh fading model. In this case, there is no correlation of phase and amplitude among the received signals. However, this assumption is not always fulfilled in practice. For instance, the transmission distance in multihop wireless transmission is supposed to be a close range and the shadowing effect needs to be taken into account. The ray launching method based on geometrical optics is a technique for estimating a deterministic propagation channel. In this paper, 3D ray launching method is used for wireless propagation prediction. By computer simulations, the performance of 2-hop cooperative relaying with two relay stations is investigated. The simulation-based performance analysis confirms that the cooperative relaying scheme has an advantage of diversity gain thus improving the bit error ratio performance.
Hiroki Tanaka, Hidekazu Murata, Koji Yamamoto 0001, Susumu Yoshida
VTC Fall1
1999 Integrated Distributed Application Management under Multi-ORB-vendor Environment
abstract
A mechanism is proposed for managing distributed applications (DAPs) in environment having multiple object request brokers (ORBs). The mechanism ensures the normal-state functioning of DAPs composed of servers. It facilitates DAP management including fault, performance, and configuration management throughout the DAP lifecycle. The mechanism includes three kinds of key components. The first is a domain manager which is responsible for fault and configuration management of DAPs. The second is a server manager, which is responsible for fault and configuration management for servers that make up DAPs. The third is a monitor that manages the performance and normal/abnormal state of servers. The common interface for monitors functioning on different kinds of DAP platforms facilitates managing DAPs over multiple ORBs. The prototype of the proposed mechanism functions properly and performs well when used with a CORBA-2.0-compliant ORE.
Hiroki Tanaka, Hidetsugu Kobayashi, Hiroshi Ishii 0002
ISADS1
1996 An environment for supporting cooperative operation over multiple networks
abstract
If a service extends over multiple networks, network operators of these networks should be able to negotiate service provision and contract establishment with each other. We have developed a support environment allowing such negotiation for a target service of network management of an ATM virtual path network. Network resource information of each network is protected against other network operators' tampering by classifying it according to the level of access permission. User friendly graphical objects in this environment facilitate the use of underlying management functions to manipulate the networks. A videoconferencing system is attached to the environment to help interaction among the operators. With all these features, this environment enables network providers to cooperate with each other to set up VP trails in quick response to customer's demands.
Tatsuo Nohara, Hiroki Tanaka, Hiroshi Ishii 0002, Osamu Miyagashi, Makoto Yamada
NOMS2