VLDB 2026 Research / reviewers in the wild / expert
Hung-Hsuan Huang
dblp:85/4098
· DBLP profile ↗
59ranked-venue papers
15as first author
3since 2021 · last 2025
0000-0002-0376-3535ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 34 · 12 first-author · 3 since 2021Artificial intelligence and machine learning · 28 · 10 first-authorDatabases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 5Graphics, computer vision, multimedia, augmented reality and games · 3Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Facial Expression Generation Model for Listening Agents Considering Speaker Engagement in Active Listening ConversationabstractThis study proposes a facial expression generation model for attentive listening behavior using conditional Generative Adversarial Networks (GAN), based on Skype-based conversations between elderly individuals and university students. The model takes as input the speaker’s non-verbal features–facial expressions and head pose and prosodic features along with the speaker’s engagement state. The output is the listener agent’s facial movements. Using Frechet Motion Distance (FMD) as the evaluation metric, we compared four models with different forms of engagement input. The results showed that using engagement as five-level categorical labels or as continuous values significantly improved the accuracy of facial expression generation. Hide Hosoda, Hung-Hsuan Huang |
HAI | 2 |
| 2024 | Do We Need To Watch It All? Efficient Job Interview Video Processing with Differentiable MaskingabstractWith technological advancements in transmitting and storing large video files, more and more organizations are incorporating asynchronous video interviews as part of their personnel selection process. Automatic evaluation of these videos is a challenging machine learning setting because the samples are composed of time series input data but only one overall label is available. It is unclear which segments of the time series input (i.e., videos) are the most important ones for prediction. Not all nonverbal features, spoken words, and utterances contribute equally to the prediction; some segments of the videos might even introduce noise to the model. Processing all multimodal information is therefore inefficient. To address this challenge, we propose a framework to model the content of the answer via the full transcription and the speaking patterns of the interviewee via short clips. Our model learns to automatically select the most informative segment by previewing the acoustic modality using a technique called differentiable masking. The results show that our method outperforms existing approaches while being more efficient since only partial multimodal data are processed, and the interpretability of the model is enhanced. Sixia Li, Candy Olivia Mawalim, Hung-Hsuan Huang, Chee Wee Leong, Shogo Okada |
ICMI | 4 |
| 2023 | Using Gamification and Multi-agent Simulation for Teacher Training in Preventing BullyingabstractThis paper presents a virtual classroom game that the trainee teacher can practice how to manage in the aspect of inter-student relationship. The virtual students are driven as a multi-agent simulation, and the player of the game can intervene with the students and try to influence their relationships. The scores are measured by the number of isolated and fringe students. Hung-Hsuan Huang |
HAI | 1 |
| 2019 | Designing a Data Corpus of Collaborative Group Tasks with the Members from Unbalanced Cultural BackgroundsabstractAccording to current estimates, the population of Japan will continuously decrease to be less than 100 million from currently 126.4 million by 2053 [1]. In order to supplement the decreasing labor force, the Japanese government has stated to relax the regulations on introducing foreignworkers. Thus, an increase of foreignerworkers in Japanese society can be expected in the near future. In such a situation, foreigner workers are the minority and have to collaboratively work with Japanese colleagues who are the majority. Our project is aiming at developing an environment for both foreigner and Japanese people to practice collaborative work in an unbalanced situation, that is, Japanese people are the majority and the task itself is Japanese favored. This paper describes our design of an experiment to acquire multimodal sensory data in collaborative tasks with unbalanced composition of cultural backgrounds. These data are supposed to be used for developing foreigner / Japanese behavior generation / detection models in a training environment with virtual agents. From our knowledge, there is no such dataset available and we believe the data collected can provide valuable resources for developing tools in supporting unbalanced groups. Kai-Yuan Ding, Hung-Hsuan Huang, Nicolas Berberich, Mineya Kaseda, Kazuhiro Kuwabara, Toyoaki Nishida |
HAI | 2 |
| 2019 | An Investigation on the Effectiveness of Multimodal Fusion and Temporal Feature Extraction in Reactive and Spontaneous Behavior Generative RNN Models for Listener AgentsabstractLike a human listener, a listener agent reacts to its communicational partners' non-verbal behaviors such as head nods, facial expressions, and voice tone. When adopting these modalities as inputs and develop the generative model of reactive and spontaneous behaviors using machine learning techniques, the issues of multimodal fusion emerge. That is, the effectiveness of different modalities, frame-wise interaction of multiple modalities, and temporal feature extraction of individual modalities. This paper describes our investigation on these issues of the task in generating of virtual listeners' reactive and spontaneous idling behaviors. The work is based on the comparison of corresponding recurrent neural network (RNN) configurations in the performance of generating listener's (the agent) head movements, gaze directions, facial expressions, and postures from the speaker's head movements, gaze directions, facial expressions, and voice tone. A data corpus recorded in a subject experiment of active listening is used as the ground truth. The results showed that video information is more effective than audio information, and frame-wise interaction of modalities is more effective than temporal characteristics of individual modalities. Hung-Hsuan Huang, Masato Fukuda, Toyoaki Nishida |
HAI | 1 |
| 2019 | Detection of Student Teacher's Intention using Multimodal Features in a Virtual ClassroomabstractThe training program for high school teachers in Japan has less opportunity to practice teaching skills. As a new practice platform, we are running a project to develop a simulation platform of school environment with computer graphics animated virtual students for students’ teachers. In order to interact with virtual students and teachers, it is necessary to estimate the intention of the teacher’s behavior and utterance. However, it is difficult to detection the teacher’s intention at the classroom only by verbal information, such as whether to ask for a response or seek a response. In this paper, we propose an automatic detection model of teacher’s intention using multimodal features including linguistic, prosodic, and gestural features. For the linguistic features, we consider the models with and without lecture contents specific information. As a result, it became clear that estimating the intention of the teacher is better when using prosodic / non-verbal information together than using only verbal information. Also, the models with contents specific information perform better. Masato Fukuda, Hung-Hsuan Huang, Toyoaki Nishida |
ICAART (1) | 2 |
| 2019 | Multimodal Assessment on Teaching Skills in a Virtual Rehearsal EnvironmentabstractIn the training programs for student teachers, the opportunity to practice teaching skills is often limited due to the lack of resources in preparing a rehearsal environment. We are developing a virtual rehearsal environment for teaching practicing with multiple virtual students. In order to provide feedbacks to the student teachers and allow them to improve their skills, automatic assessment on their performance is required. However, it is hard to assess on teaching because the assessment is subjective and often depends on the tacit knowledge of experienced teacher trainers. In this work, we proposed an automatic assessment model based on human assessment done by experienced high school teachers. Masato Fukuda, Hung-Hsuan Huang, Kazuhiro Kuwabara, Toyoaki Nishida |
IVA | 2 |
| 2019 | Development of a Platform for RNN Driven Multimodal Interaction with Embodied Conversational AgentsabstractThis paper describes our ongoing project to build a platform that enables real-time multimodal interaction with embodied conversational agents. All of the components are in modular design and can be switched to other models easily. A prototype listener agent has been developed upon the platform. Its spontaneous reactive behaviors are trained from a multimodal data corpus collected in a human-human conversation experiment. Two Gated Recurrent Unit (GRU) based models are switched when the agent is speaking or is not speaking. These models generate the agent's facial expressions, head movements, and postures from the corresponding behaviors of the human user in real-time. Benefits from the flexible design, the utterance generation part can be an autonomous dialogue manager with hand crafted rules, an on-line chatbot engine, or a human operator. Hung-Hsuan Huang, Masato Fukuda, Toyoaki Nishida |
IVA | 1 |
| 2018 | Conversation Strategy of a Chatbot for Interactive Recommendations
Yuichiro Ikemoto, Varit Asawavetvutt, Kazuhiro Kuwabara, Hung-Hsuan Huang |
ACIIDS (1) | 4 |
| 2018 | Investigation of Class Atmosphere Cognition in a VR ClassroomabstractOur aim is to develop a VR virtual classroom environment that can interact with virtual students through atmosphere. The trainee im- proves his / her class while confirming the change of atmosphere of the virtual students. For such a virtual environment that can give feedback autonomously expressing the atmosphere, data on the atmosphere that can be perceived as feedback are required. In this research, we conducted experiments to verify whether the at- mosphere exuded by virtual students, multi agents, is perceivable through a VR environment. Masato Fukuda, Hung-Hsuan Huang, Toyoaki Nishida |
HAI | 2 |
| 2018 | Integration of DNN Generated Spontaneous Reactions with a Generic Multimodal Framework for Embodied Conversational AgentsabstractThis paper describes the recent extensions of our previously proposed GECA Framework in incorporating the interconnection with GECA original networking library and ZeroMQ message passing library, the connection with Python based DNN developing API, Keras, and a character animator supporting FACS based detailed facial animations and lip-syncing with real-time generated voice tracks. Hung-Hsuan Huang, Masato Fukuda, Stef van der Struijk, Toyoaki Nishida |
HAI | 1 |
| 2018 | Proposal of a Multi-purpose and Modular Virtual Classroom Framework for Teacher TrainingabstractThis paper presents our proposal of platform with virtual classroom and virtual students which allows to build a flexible system such as autonomous system and evaluation system with combination of modules. Masato Fukuda, Hung-Hsuan Huang, Kazuhiro Kuwabara, Toyoaki Nishida |
IVA | 2 |
| 2018 | FACSvatar: An Open Source Modular Framework for Real-Time FACS based Facial AnimationabstractEmbodied Conversational Agents often employ advanced multimodal analysis of human users for affective inference, however, its facial expressions in response are often pre-made or coded animations. Generated data-driven facial animations have the advantage that they can be more natural and do not require a database at run-time. The open source modular framework FACSvatar is presented, which processes and animates FACS based data in real-time. Tools that create 3D human models are supported and facial data can be visualized in popular tools in the gaming industry. All functionality is split-up into modules set-up in a publisher-subscriber pattern to provide easy integration with other platforms. A deep learning module for real-time generation of AU data for data-driven animation is implemented. A user evaluation of their expressions being animated through our framework was done. Their ratings were slightly positive, but more improvements have to be made in terms of data quality and individual fine-tuning. Also, the modules' latency and performance have been measured. On average, FACS data is visualized in 28.55 ms. Stef van der Struijk, Hung-Hsuan Huang, Maryam Sadat Mirzaei, Toyoaki Nishida |
IVA | 2 |
| 2017 | International Business Matching Using Word Embedding
Didier Gohourou, Daiki Kurita, Kazuhiro Kuwabara, Hung-Hsuan Huang |
ACIIDS (1) | 4 |
| 2017 | Proposal of a Parameterized Atmosphere Generation Model in a Virtual ClassroomabstractThe training program of high school teachers in Japan lacks the chance to practice teaching skills and the admission of classes. The result is, many young teachers left their jobs in the first year due to frustration and other mental issues. In order to relieve this issue, we are running a project to develop a platform of virtual school environment and virtual students which allows the teacher trainees to practice with. As part of this platform, in order to send feedbacks to the trainee to stimulate them to improve their performance, we propose the use of the whole group of virtual students to generate per- ceivable atmospheres. Three research issues then emerge:(1) Whether the atmospheres in a classroom can be expressed in a computational (parameterized) model? (2) If the answer to (1) is 'yes', what are the elements to construct atmospheres? (3) The perception of atmospheres can be considered as very subjective, how to make the model objective becomes the last research issue. This paper presents our investigation on these issues and propose a parameterized atmosphere generation model based on empirical results. Masato Fukuda, Hung-Hsuan Huang, Naoki Ohta, Kazuhiro Kuwabara |
HAI | 2 |
| 2017 | Proposal of a Model to Determine the Attention Target for an Agent in Group Discussion with Non-verbal FeaturesabstractIn recent years, companies are seeking for communication skill from their employers. More and more companies adopt group discussions in employer recruitment to evaluate the ap- plicants' communication skill. However, the opportunity to improve communication skill in group discussion is limited due to the lack of partners. In order to solve this issue, our ongoing project is aiming to build a virtual agent or a robot that can participate group discussion, so that its users can re- peatedly practice group discussion with it. In this paper, we propose the models in directing the agent's attention toward the other participants in three situations:when the agent is speaking, when the agent is listening, and when no partic- ipant is speaking. First, we gathered a data corpus of the discussion of 10 four-people groups. We then use low-level non-verbal features including attention of other participant, voice prosody, head movements, and speech turn extracted in the 10-hour corpus to train support vector machine models to determine the agent's attention on the other participants, or the material. The performance of the detection models in F-measure range between 0.4 and 0.6. Seiya Kimura, Hung-Hsuan Huang, Shogo Okada, Naoki Ohta, Kazuhiro Kuwabara |
HAI | 2 |
| 2017 | Toward Gamified Knowledge Contents Refinement - Case Study of a Conversation Partner Agent
Takayuki Iwamae, Kazuhiro Kuwabara, Hung-Hsuan Huang |
ICAART (1) | 3 |
| 2016 | Knowledge Base Refinement with Gamified Crowdsourcing
Daiki Kurita, Boonsita Roengsamut, Kazuhiro Kuwabara, Hung-Hsuan Huang |
ACIIDS (1) | 4 |
| 2016 | Development of a Simulated Environment for Recruitment Examination and Training of High School TeachersabstractWhile the environment of schools become more and more complicated, the improvement of teachers' skills in teaching and management is required. In this study, we focus on the development of a Wizard-of-Oz (WOZ) platform of simulated school environment, which can be utilized for teacher training or the examination of teacher recruitment from remote. This system is comprised of two front ends, one is a simulated classroom for the trainee, the other one is the interface for the system operator / investigator. The virtual classroom contains a number of virtual students who are controlled by the operator from the remote. The operator can observe the trainee from a dedicated interface and control the behaviors of any individual student as well as the atmosphere of the whole class. The whole-class atmosphere created by relatively large number of students is modeled as a concentration-arousal two dimensional space. The prototype system is evaluated with subject experiment and the results are reported. Masato Fukuda, Hung-Hsuan Huang, Tetsuya Kanno, Naoki Ohta, Kazuhiro Kuwabara |
HAI | 2 |
| 2016 | Toward a Guide Agent who Actively Intervene Inter-user Conversation - Timing Definition and Trial of Automatic Detection using Low-level Nonverbal FeaturesabstractAs the advance of embodied conversational agent (ECA) technologies, there are more and more real-world deployed applications of ECAâ??s. The guides in museums or exhibitions are typical examples. However, in these situations, the agent systems usually need to engage groups of visitors rather than individual ones. In such a multi-user situation, which is much more complex than single user one, specialized additional features are required. One of them is the ability for the agent to smoothly intervene user-user conversation. In order to realize this, at first, a Wizard-of-Oz (WOZ) experiment was conducted for collecting human interaction data. By analyzing the collected data corpus, four kinds of timings that potentially allow the agent to do intervention were found. The collected corpus was then annotated with these defined timings by recruited evaluators with a dedicated and intuitive tool. Finally, as the trial of the possibility of automatic detection on these timings, the use of non-verbal low level features were able to achieve a moderate accuracy. Hung-Hsuan Huang, Shochi Otogi, Ryo Hotta, Kyoji Kawagoe |
ICAART (2) | 1 |
| 2016 | Estimating communication skills using dialogue acts and nonverbal features in multiple discussion datasetsabstractThis paper focuses on the computational analysis of the individual communication skills of participants in a group. The computational analysis was conducted using three novel aspects to tackle the problem. First, we extracted features from dialogue (dialog) act labels capturing how each participant communicates with the others. Second, the communication skills of each participant were assessed by 21 external raters with experience in human resource management to obtain reliable skill scores for each of the participants. Third, we used the MATRICS corpus, which includes three types of group discussion datasets to analyze the influence of situational variability regarding to the discussion types. We developed a regression model to infer the score for communication skill using multimodal features including linguistic and nonverbal features: prosodic, speaking turn, and head activity. The experimental results show that the multimodal fusing model with feature selection achieved the best accuracy, 0.74 in R2 of the communication skill. A feature analysis of the models revealed the task-dependent and task-independent features to contribute to the prediction performance. Shogo Okada, Yoshihiko Ohtake, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Yutaka Takase, Katsumi Nitta |
ICMI | 5 |
| 2016 | Composing the Atmosphere of a Virtual Classroom with a Group of Student Agents
Masato Fukuda, Hung-Hsuan Huang, Naoki Ohta, Kazuhiro Kuwabara |
IVA | 2 |
| 2016 | Development of a Virtual Classroom for High School Teacher Training
Hung-Hsuan Huang, Yuki Ida, Kohei Yamaguchi, Kyoji Kawagoe |
IVA | 1 |
| 2016 | Toward a Conversation Partner Agent for People with Aphasia: Assisting Word Retrieval
Kazuhiro Kuwabara, Takayuki Iwamae, Yudai Wada, Hung-Hsuan Huang, Keisuke Takenaka |
KES-IDT (1) | 4 |
| 2016 | Introduction to the Special Issue on New Directions in Eye Gaze for Interactive Intelligent SystemsabstractEye gaze has been used broadly in interactive intelligent systems. The research area has grown in recent years to cover emerging topics that go beyond the traditional focus on interaction between a single user and an interactive system. This special issue presents five articles that explore new directions of gaze-based interactive intelligent systems, ranging from communication robots in dyadic and multiparty conversations to a driving simulator that uses eye gaze evidence to critique learners’ behavior. Yukiko I. Nakano, Roman Bednarik, Hung-Hsuan Huang, Kristiina Jokinen |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Searching human actions based on a multi-dimensional time series similarity calculation methodabstractWith the rapid performance improvement and popularization of sensor devices, a large amount of human action data can be captured in databases. Classification, recognition, searching, and mining of such human actions are promising applications. Although many of these applications have been developed, searching the large quantity of data, especially given the high dimensionality of the captured temporal data sequence is time-consuming. To reduce this time cost, we use a novel method for approximating a multi-dimensional time-series, named multi-dimensional time-series Approximation with use of Local features at Thinned-out Keypoints (A-LTK). With A-LTK applications for two human motion types, sign language and dancing, we found that the categorization of human action data and the search for the most similar human action became more accurate and reduced the time cost. Kosuke Sugano, Kenta Oku, Hung-Hsuan Huang, Kyoji Kawagoe |
ICIS | 4 |
| 2015 | SQUED: A novel crowd-sourced system for detection and localization of unexpected events from smartphone-sensor dataabstractIn this paper, a Smart and Quick Unexpected-Event Detector, SQUED, is proposed to detect and localize unexpected events, such as traffic accidents, using crowd sourcing with smartphone devices. When users find an event, they point their smartphones in the direction of the event. SQUED can determine the location and the time of the event using global positioning system (GPS) and geomagnetic sensor data transferred from the built-in devices. In SQUED, a user doesn't need to swipe or tap; he simply points the device. We developed SQUED on Android smartphones and evaluated our event-detection method in real-world settings. It was shown that SQUED can stably detect an event with high precision, even with noisy and fluctuating data for observed locations and directions. Taishi Yamamoto, Kenta Oku, Hung-Hsuan Huang, Kyoji Kawagoe |
ICIS | 3 |
| 2015 | A human behavior processes database prototype system for surgery supportabstractDespite the rapid progress in the development of sensor technologies, as well as of information management using database technology, no appropriate database system exists for comprehensively archiving and searching human behavior processes. In this paper, we propose a human behavior database prototype system which can archive and support surgery activities. The database system can be used to find human processes similar to a given process and distinguish the differences between them. The significance of this prototype system is that it can implement a comprehensive archive and an efficient search of complex human activities. The system can be used to support medical skill training by presenting users with the proper surgery processes, and the differences that need to be noticed. In this paper, we first describe the basic concepts and structure design of the prototype system. We also discuss how our proposed system is suitable for human behavior process management. Zhang Zuo, Kenta Oku, Hung-Hsuan Huang, Kyoji Kawagoe |
ICIS | 3 |
| 2015 | Toward gamification of knowledge base constructionabstractKnowledge is integral to an intelligent system. However, constructing a knowledge base often involves great cost and substantial effort on the part of human experts in the target domain. In this paper, we present an approach that harnesses the power of casual users to construct the knowledge base. The users answer a simple series of questions and, by aggregating their answers, we obtain the information necessary to augment and refine the knowledge base. To motivate the users to engage in the process, we presented the questions via a gamification approach. Our experiments with the knowledge base of an Apartment FAQ system substantiate the potential of our proposed approach. Boonsita Roengsamut, Kazuhiro Kuwabara, Hung-Hsuan Huang |
INISTA | 3 |
| 2014 | POSTER: Security Control System Enabling to Keep an Intra-LAN in a Secure State Using Security-and-Performance Ratio Control PoliciesabstractWith the emergence of inexpensive network components and high-speed network services, a variety of network-capable electronic devices have become available. However, as no self-defence mechanism is equipped in a usual network device, the device can be attacked by other devices or by attackers, which causes it in a vulnerable state. In this paper, we propose a system for keeping an Intra-LAN in a secure state with isolation of such a network information device. The system the communications between devices in the LAN by introducing the safety-and-performance ratio control policies. The control policies are controlled dynamically. The isolation of a devices is performed using an intrusion detection system. A secure LAN can be maintained by the dynamic controlled policies.We have confirmed that the LAN can keep secure with actual LAN environment experiments of our system. Yutaka Juba, Hung-Hsuan Huang, Kyoji Kawagoe |
CCS | 2 |
| 2014 | Analysis of personality traits for intervention scene detection in multi-user conversationabstractAs the advance of embodied conversational agent (ECA) technologies, there are more and more real-world deployed applications of ECA's like the guides in museums or exhibitions. However, in these applications, the agent systems are usually used by groups of visitors rather than individuals. In such multi-user situation, which is more complex sophisticated than single user one, specific features are required. There can be difference in how and when to intervene in the conversation of others due to the variety of personality. In order to realize a more implement the human-like and more helpful guide agent, this work tries to explore the relationship between personality and the willing to intervene in users' conversation as the role of a guide. In this paper, the analysis results of the intervention action and personality traits are reported. Shochi Otogi, Hung-Hsuan Huang, Ryo Hotta, Kyoji Kawagoe |
HAI | 2 |
| 2014 | Gaze-in 2014: the 7th Workshop on Eye Gaze in Intelligent Human Machine InteractionabstractThis paper presents a summary of the seventh workshop on Eye Gaze in Intelligent Human Machine Interaction. The Gaze-in 2014 workshop is a part of a series of workshops held around the topics related to gaze and multimodal interaction. The workshop web-site can be found at http://hhhuang.homelinux.com/gaze_in/. Hung-Hsuan Huang, Roman Bednarik, Kristiina Jokinen, Yukiko I. Nakano |
ICMI | 1 |
| 2014 | Predicting Influential Statements in Group Discussions using Speech and Head Motion InformationabstractGroup discussions are used widely when generating new ideas and forming decisions as a group. Therefore, it is assumed that giving social influence to other members through facilitating the discussion is an important part of discussion skill. This study focuses on influential statements that affect discussion flow and highly related to facilitation, and aims to establish a model that predicts influential statements in group discussions. First, we collected a multimodal corpus using different group discussion tasks; in-basket and case-study. Based on schemes for analyzing arguments, each utterance was annotated as being influential or not. Then, we created classification models for predicting influential utterances using prosodic features as well as attention and head motion information from the speaker and other members of the group. In our model evaluation, we discovered that the assessment of each participant in terms of discussion facilitation skills by experienced observers correlated highly to the number of influential utterances by a given participant. This suggests that the proposed model can predict influential statements with considerable accuracy, and the prediction results can be a good predictor of facilitators in group discussions. Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, Shogo Okada |
ICMI | 4 |
| 2014 | Music classification method based on lyrics for music therapyabstractMusic is used for people practicing sports, for elderly individuals, and to help train the mind. Recently in music information science, studies have been conducted on music therapy and on music classification from a therapeutic point of view. However, most of these studies have classified music based on melody and tempo. No classification method that is based on lyrics has been performed for music therapy support. In this paper, we propose a music classification method using emotional words included in lyrics toward music therapy. Music is classified by using a clustering method. In addition, this paper shows that a higher classification precision is obtained when compared with the conventional Mizuki Furuya, Hung-Hsuan Huang, Kyoji Kawagoe |
IDEAS | 2 |
| 2014 | Exploring the Difference of the Impression on Human and Agent Listeners in Active Listening Dialog
Hung-Hsuan Huang, Natsumi Konishi, Sayumi Shibusawa, Kyoji Kawagoe |
IVA | 1 |
| 2014 | GUIDES: a graphical user identifier scheme using sketching for mobile web-servicesabstractIn this paper, a novel graphical user identifier scheme for web-based mobile services. Despite a rapid progress in biological authentication technologies for user identification, the traditional text-based UserID-and-password scheme has still widely been used even for mobile web-services. In the mobile usage, a virtual keyboard is difficult to use for input texts, which causes much time to input a text. To decrease the difficulty, GUIDES (Graphical User IDEntifier using Sketching) is proposed in this paper. With GUIDES, a user can input its user identifier, called GUID (Graphical User ID), by drawing a sketch on a mobile touchscreen device. From our experiments, it is concluded that GUIDES can enable a user input his/her User-ID more efficiently and remember it more, compared with the text-based user-id scheme. Yusuke Matsuno, Hung-Hsuan Huang, Kyoji Kawagoe |
Mobile HCI | 2 |
| 2013 | Towards Music Information Retrieval driven by EEG signals: Architecture and preliminary experimentsabstractAlthough much research on Music Information Retrieval (MIR) has been done in the last decade, the input of the current MIR to specify a user query for finding a similar piece of music is still either by the existing old-fashioned keywords or by music contents. We aim to realize a new type of MIR equipped with brain-computer interfaces using electroencephalogram (EEG) signals. Toward the new MIR, we propose an architecture of MIR driven by EEG signals in this paper. While the architecture contains many issues to be solved, the point of the architecture is to construct user's music query in multi-layered aggregation of EEG signals. We describe in this paper the preliminary experiments conducted for selecting some appropriate low-level features for our multi-layered query construction and matching. It is obtained that the mental states of users while listening to music can be classified with high accuracy by using EEG signal aggregated features. We are starting development of detailed design of the architecture using the results described in the paper. Yuuko Morita, Hung-Hsuan Huang, Kyoji Kawagoe |
ICIS | 2 |
| 2013 | Graphical password using object-based image rankingabstractIn this paper, we propose a new graphical password using object-based image ranking, called OBIR, which enables appropriate images to be presented to users during authentication. Research on graphical password is being conducted and receiving more and more public attention due to its potential of being an alternative for textual password. However, the main problem of graphical password is its vulnerability to shoulder surfing attacks, especially on mobile devices where the login password is easily visible in public. In order to overcome this issue, we propose a novel graphical password using image ranking method based on the objects in the image itself. The higher the ranking of an image is, the more appropriate it is to be user's selection without exposing too much information of the password to the shoulder-surfer. Upon our experiments, it is obtained that our image ranking method is effective in filtering appropriate pass-image that even if it is selected in public, the password is safe and therefore resistant to shoulder surfing. Cuong Xuan Nguyen, Hung-Hsuan Huang, Kyoji Kawagoe |
CCS | 2 |
| 2013 | Gazein'13: the 6th workshop on eye gaze in intelligent human machine interaction: gaze in multimodal interactionabstractThis paper presents a summary of the sixth workshop in Eye Gaze in Intelligent Human Machine Interaction. The GazeIn'13 workshop is a part of a series of workshops held around the topics related to gaze and multimodal interaction. Roman Bednarik, Hung-Hsuan Huang, Yukiko I. Nakano, Kristiina Jokinen |
ICMI | 2 |
| 2013 | Implementation and evaluation of a multimodal addressee identification mechanism for multiparty conversation systemsabstractIn conversational agents with multiparty communication functionality, a system needs to be able to identify the addressee for the current floor and respond to the user when the utterance is addressed to the agent. This study proposes some addressee identification models based on speech and gaze information, and tests whether the models can be applied to different proxemics. We build an addressee identification mechanism by implementing the models and incorporate it into a fully autonomous multiparty conversational agent. The system identifies the addressee from online multimodal data and uses this information in language understanding and dialogue management. Finally, an evaluation experiment shows that the proposed addressee identification mechanism works well in a real-time system, with an F-measure for addressee estimation of 0.8 for agent-addressed utterances. We also found that our system more successfully avoided disturbing the conversation by mistakenly taking a turn when the agent is not addressed. Yukiko I. Nakano, Naoya Baba, Hung-Hsuan Huang, Yuki Hayashi |
ICMI | 3 |
| 2013 | Detection of Most Popular Routes and Effective Time Segments Using Trajectory DistributionsabstractThere have been some innovative research studies for detecting the Most Popular Route (MRP) using GPS devices in order to support tourists who travel in an unfamiliar area. The MPR is a route on which many moving objects move the most among the entire possible routes. Current MRP detection methods do not take into account the time of trajectory measurement, however, road conditions vary depending on a time zone. Therefore, the detected MRP may not be an appropriate route that was defined outside of the certain time zone. The aim of this study is to propose a new method to detect the MRP which is capable of considering a time zone of trajectory measurement. In addition to the new method, "Popularity Measure" is proposed in order to verify the suitability of the detected MRP. The detected MRP using the existing and proposed method are evaluated by compared from a viewpoint of this popularity measures. Kazuma Ito, Hung-Hsuan Huang, Kyoji Kawagoe |
ISM | 2 |
| 2013 | Detecting Musical Genre Borders for Multi-label Genre ClassificationabstractIn this paper, we propose a novel method to detect music genre borders for the music genre classification. The music genre classification is getting more important because music is influenced by an increasing amount of different musical styles. A general approach to classify music genres is a single genre labeling that usually gives the meaning of inherent stylistic elements to a musical piece. However this gives ambiguity in case of a music piece having multiple genres. To solve the problem, we consider separating the multi-label classification task into the single-label genre classification task. We propose a novel method to detect music genre borders for multi-label genre classification. The proposed method can find borderlines of different genres in music. Moreover, it is strongly expected to realize the multi-label genre classification to apply the single-label genre classification to each detected music segment. Hiroki Nakamura, Hung-Hsuan Huang, Kyoji Kawagoe |
ISM | 2 |
| 2013 | Dynamic Isolation of Network Devices Using OpenFlow for Keeping LAN Secure from Intra-LAN AttackabstractWith the emergence of inexpensive network components and high-speed network services, a variety of network-capable electronic devices have become available. Typical examples of such devices include printers, network access storage (NAS), and video recorders. Because software on these devices is not always kept up-to-date, the devices are susceptible to intra-local- area-network (LAN) attacks. In this paper, a novel network system architecture is proposed to protect network devices from intra-LAN attacks by dynamically isolating infected devices with OpenFlow. Preliminary evaluation results demonstrate that the architecture is effective in actual LAN environments. Yutaka Juba, Hung-Hsuan Huang, Kyoji Kawagoe |
KES | 2 |
| 2012 | Predicting Online Auction Final Prices Using Time Series Splitting and Clustering
Takuya Yokotani, Hung-Hsuan Huang, Kyoji Kawagoe |
APWeb | 2 |
| 2012 | Tag Association Based Graphical Password Using Image Feature Matching
Kyoji Kawagoe, Shinichi Sakaguchi, Yuki Sakon, Hung-Hsuan Huang |
DASFAA (2) | 4 |
| 2012 | 4th workshop on eye gaze in intelligent human machine interaction: eye gaze and multimodalityabstractThis is the fourth workshop in a series of workshops on Eye Gaze in Intelligent Human Machine Interaction, in which we have discussed a wide range of issues for eye gaze; technologies for sensing human attentional behaviors, roles of attentional behaviors as social gaze in human-human and human-humanoid interaction, attentional behaviors in problem-solving and task-performing, gaze-based intelligent user interfaces, and evaluation of gaze-based user interfaces. In addition to these topics, this year's workshop focuses on eye gaze in multimodal interpretation and generation. Since eye gaze is one of the facial communication modalities, gaze information can be combined with other modalities or bodily motions to contribute to the meaning of utterance and serve as communication signals. Yukiko I. Nakano, Kristiina Jokinen, Hung-Hsuan Huang |
ICMI | 3 |
| 2012 | Photo-Taking Point Recommendation with Nested ClusteringabstractIn this paper, we propose a novel recommendation method for photo-taking points from a large amount of social community photo collections. There are many research activities on photo-related recommendations from a lot of photos stored and managed by photo sharing web services, such as Flickr, Picas a and Panoramio, Although some methods, such as landmark recommendation, tag recommendation and photo recommendation have already been proposed, no photo-taking point recommendation methods have been realized yet for social photo collections. In order to realize photo-taking point recommendation, we introduce a novel point and photo selection method based on nested clustering. From our experiments, it is shown that better recommendation accuracy with our proposed method can be attained. Kosuke Kimura, Hung-Hsuan Huang, Kyoji Kawagoe |
ISM | 2 |
| 2012 | Modeling the Multi-modal Behaviors of a Virtual Instructor in Tutoring Ballroom Dance
Hung-Hsuan Huang, Yuki Seki, Masaki Uejou, Joo-Ho Lee 0001, Kyoji Kawagoe |
IVA | 1 |
| 2011 | Making virtual conversational agent aware of the addressee of users' utterances in multi-user conversation using nonverbal informationabstractIn multi-user human-agent interaction, the agent should respond to the user when an utterance is addressed to it. To do this, the agent needs to be able to judge whether the utterance is addressed to the agent or to another user. This study proposes a method for estimating the addressee based on the prosodic features of the user's speech and head direction (approximate gaze direction). First, a WOZ experiment is conducted to collect a corpus of human-humanagent triadic conversations. Then, analysis is performed to find out whether the prosodic features as well as head direction information are correlated with the addressee-hood. Based on this analysis, a SVM classifier is trained to estimate the addressee by integrating both the prosodic features and head movement information. Finally, a prototype agent equipped with this real-time addressee estimation mechanism is developed and evaluated. Hung-Hsuan Huang, Naoya Baba, Yukiko I. Nakano |
ICMI | 1 |
| 2011 | Identifying Utterances Addressed to an Agent in Multiparty Human-Agent Conversations
Naoya Baba, Hung-Hsuan Huang, Yukiko I. Nakano |
IVA | 2 |
| 2011 | Source Orientation in Communication with a Conversational Agent
Yugo Hayashi, Hung-Hsuan Huang, Victor V. Kryssanov, Akira Urao, Kazuhisa Miwa, Hitoshi Ogawa |
IVA | 2 |
| 2011 | Toward a Conversational Virtual Instructor of Ballroom Dance
Masaki Uejou, Hung-Hsuan Huang, Joo-Ho Lee 0001, Kyoji Kawagoe |
IVA | 2 |
| 2009 | The Lessons Learned in Developing Multi-user Attentive Quiz Agents
Hung-Hsuan Huang, Takuya Furukawa, Hiroki Ohashi, Aleksandra Cerekovic, Yuji Yamaoka, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 1 |
| 2007 | Towards a Multicultural ECA Tour Guide System
Aleksandra Cerekovic, Hung-Hsuan Huang, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 2 |
| 2007 | A Script Driven Multimodal Embodied Conversational Agent Based on a Generic Framework
Hung-Hsuan Huang, Aleksandra Cerekovic, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 1 |
| 2007 | A Quiz Game Console Based on a Generic Embodied Conversational Agent Framework
Hung-Hsuan Huang, Taku Inoue, Aleksandra Cerekovic, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
IVA | 1 |
| 2006 | Toward a Universal Platform for Integrating Embodied Conversational Agent Components
Hung-Hsuan Huang, Tsuyoshi Masuda, Aleksandra Cerekovic, Kateryna Tarasenko, Igor S. Pandzic, Yukiko I. Nakano, Toyoaki Nishida |
KES (2) | 1 |
| 2005 | Generating CG Movies Based on a Cognitive Model of Shot Transition
Kazunori Okamoto, Yukiko I. Nakano, Masashi Okamoto, Hung-Hsuan Huang, Toyoaki Nishida |
KES (3) | 4 |
| 2004 | Gallery: In Support of Human Memory
Hung-Hsuan Huang, Yasuyuki Sumi, Toyoaki Nishida |
KES | 1 |