EDBT 2026 Demo / reviewers in the wild / expert
Shinya Fujie
dblp:14/4942
· DBLP profile ↗
24ranked-venue papers
5as first author
8since 2021 · last 2023
0009-0007-5469-9207ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Conversation-Oriented ASR with Multi-Look-Ahead CBS ArchitectureabstractDuring conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognition (ASR) used for transcribing the speech in real-time must achieve high accuracy without delay. In streaming ASR, high accuracy is assured by attending to look-ahead frames, which leads to delay increments. To tackle this trade-off issue, we propose a multiple latency streaming ASR to achieve high accuracy with zero look-ahead. The proposed system contains two encoders that operate in parallel, where a primary encoder generates accurate outputs utilizing look-ahead frames, and the auxiliary encoder recognizes the look-ahead portion of the primary encoder without look-ahead. The proposed system is constructed based on contextual block streaming (CBS) architecture, which leverages block processing and has a high affinity for the multiple latency architecture. Various methods are also studied for architecting the system, including shifting the network to perform as different encoders; as well as generating both encoders’ outputs in one encoding pass. Huaibo Zhao, Shinya Fujie, Tetsuji Ogawa, Jin Sakuma, Yusuke Kida, Tetsunori Kobayashi |
ICASSP | 2 |
| 2023 | Multimodal Turn-Taking Model Using Visual Cues for End-of-Utterance Prediction in Spoken Dialogue Systems
Fuma Kurata, Mao Saeki, Shinya Fujie, Yoichi Matsuyama |
INTERSPEECH | 3 |
| 2023 | Improving the response timing estimation for spoken dialogue systems by reducing the effect of speech recognition delay
Jin Sakuma, Shinya Fujie, Huaibo Zhao, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2022 | Confusion Detection for Adaptive Conversational Strategies of An Oral Proficiency Assessment Interview Agent
Mao Saeki, Kotoka Miyagi, Shinya Fujie, Shungo Suzuki, Tetsuji Ogawa, Tetsunori Kobayashi, Yoichi Matsuyama |
INTERSPEECH | 3 |
| 2022 | Response Timing Estimation for Spoken Dialog System using Dialog Act Estimation
Jin Sakuma, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2022 | Response Timing Estimation for Spoken Dialog Systems Based on Syntactic Completeness PredictionabstractAppropriate response timing is very important for achieving smooth dialog progression. Conventionally, prosodic, temporal and linguistic features have been used to determine timing. In addition to the conventional parameters, we propose to utilize the syntactic completeness after a certain time, which represents whether the other party is about to finish speaking. We generate the next token sequence from intermediate speech recognition results using a language model and obtain the probability of the end of utterance appearing$K$tokens ahead, where$K$varies from 1 to$M$. We obtain an$M$-dimensional vector, which we denote as estimates of syntactic completeness (ESC). We evaluated this method on a simulated dialog database of a restaurant information center. The results confirmed that considering ESC improves the performance of response timing estimation, especially the accuracy in quick responses, compared with the method using only conventional features. Jin Sakuma, Shinya Fujie, Tetsunori Kobayashi |
SLT | 2 |
| 2021 | Timing Generating Networks: Neural Network Based Precise Turn-Taking Timing Prediction in Multiparty Conversation
Shinya Fujie, Hayato Katayama, Jin Sakuma, Tetsunori Kobayashi |
Interspeech | 1 |
| 2021 | Personalized Extractive Summarization for a News Dialogue SystemabstractIn modern society, people's interests and preferences are diversifying. Along with this, the demand for personalized summarization technology is increasing. In this study, we propose a method for generating summaries tailored to each user's interests using profile features obtained from questionnaires administered to users of our spoken-dialogue news delivery system. We propose a method that collects and uses the obtained user profile features to generate a summary tailored to each user's interests, specifically, the sentence features obtained by BERT and user profile features obtained from the questionnaire result. In addition, we propose a method for extracting sentences by solving an integer linear programming problem that considers redundancy and context coherence, using the degree of interest in sentences estimated by the model. The results of our experiments confirmed that summaries generated based on the degree of interest in sentences estimated using user profile information can transmit information more efficiently than summaries based solely on the importance of sentences. Hiroaki Takatsu, Mayu Okuda, Yoichi Matsuyama, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
SLT | 5 |
| 2019 | Recognition of Intentions of Users' Short Responses for Conversational News Delivery System
Hiroaki Takatsu, Katsuya Yokoyama, Yoichi Matsuyama, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 5 |
| 2018 | Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users' Interest Level
Masahiro Araki, Sayaka Tomimasu, Mikio Nakano, Kazunori Komatani, Shogo Okada, Shinya Fujie, Hiroaki Sugiyama |
LREC | 6 |
| 2018 | Investigation of Users' Short Responses in Actual Conversation System and Automatic Recognition of their IntentionsabstractIn human-human conversations, listeners often convey intentions to speakers through feedback consisting of reflexive short responses. The speakers recognize these intentions and change the conversational plans to make communication more efficient. These functions are expected to be effective in human-system conversations also; however, there is only a few systems using these functions or a research corpus including such functions. We created a corpus that consists of users' short responses to an actual conversation system and developed a model for recognizing the intention of these responses. First, we categorized the intention of feedback that affects the progress of conversations. We then collected 15604 short responses of users from 2060 conversation sessions using our news-delivery conversation system. Twelve annotators labeled each utterance based on intention through a listening test. We then designed our deep-neural-network-based intention recognition model using the collected data. We found that feedback in the form of questions, which is the most frequently occurring expression, was correctly recognized and contributed to the efficiency of the conversation system. Katsuya Yokoyama, Hiroaki Takatsu, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
SLT | 4 |
| 2016 | A Spoken Dialog System for Coordinating Information Consumption and ExplorationabstractPassive consumption of information is boring in most cases and even painful in some cases, especially when the information content is delivered by employing speech media. The user of a speech-based information delivery system, for example a text-to-speech system, usually cannot interrupt the ongoing information flow, inhibiting her/him to confirm some part of the content, or to pose an inquiry for further information exploration. We argue that a carefully designed spoken dialog system could remedy these undesirable situations, and further enable an enjoyable conversation with the users. The key technologies to realize such an attractive dialog system are: (1) pre-compilation of a dialog plan based on the analysis of a source content, and (2) the dynamic recognition of user's state of understanding and interests. This paper illustrates technical views to implement these functionalities, and discusses a dialog example to exemplify the technical merits of the proposed system. Shinya Fujie, Ishin Fukuoka, Asumi Mugita, Hiroaki Takatsu, Yoshihiko Hayashi, Tetsunori Kobayashi |
CHIIR | 1 |
| 2015 | Four-participant group conversation: A facilitation robot controlling engagement density as the fourth participantabstractIn this paper, we present a framework for facilitation robots that regulate imbalanced engagement density in a four-participant conversation as the forth participant with proper procedures for obtaining initiatives. Four is the special number in multiparty conversations. In three-participant conversations, the minimum unit for multiparty conversations, social imbalance, in which a participant is left behind in the current conversation, sometimes occurs. In such scenarios, a conversational robot has the potential to objectively observe and control situations as the fourth participant. Consequently, we present model procedures for obtaining conversational initiatives in incremental steps to harmonize such four-participant conversations. During the procedures, a facilitator must be aware of both the presence of dominant participants leading the current conversation and the status of any participant that is left behind. We model and optimize these situations and procedures as a partially observable Markov decision process (POMDP), which is suitable for real-world sequential decision processes. The results of experiments conducted to evaluate the proposed procedures show evidence of their acceptability and feeling of groupness. Yoichi Matsuyama, Iwao Akiba, Shinya Fujie, Tetsunori Kobayashi |
Comput. Speech Lang. | 3 |
| 2015 | Automatic Expressive Opinion Sentence Generation for Enjoyable Conversational SystemsabstractIn terms of functional conversations, Grice's Maxim of Quantity suggests that responses should contain no more information than was explicitly asked for. However, in our daily conversations, more informative response skills are usually employed in order to hold enjoyable conversations with interlocutors. These responses are usually produced as forms of one's additional opinions, which usually contain their original viewpoints as well as novel means of expression, rather than simple and common responses characteristic of the general public. In this paper, we propose automatic expressive opinion sentence generation mechanisms for enjoyable conversational systems. The generated opinions are extracted from a large number of reviews on the web, and ranked in terms of contextual relevance, length of sentences, and amount of information represented by the frequency of adjectives. The sentence generator also has an additional phrasing skill. Three controlled lab experiments were conducted, where subjects were requested to read generated sentences and watch videos filmed about conversations between the robot and a person. The results implied that mechanisms effectively promote users' enjoyment and interests. Yoichi Matsuyama, Akihiro Saito, Shinya Fujie, Tetsunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | Psychological evaluation of a group communication activation robot in a party game
Yoichi Matsuyama, Shinya Fujie, Hikaru Taniyama, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2009 | System design of group communication activator: an entertainment task for elderly careabstractOur community is facing serious Aging Society especially in Japan. Yoichi Matsuyama, Hikaru Taniyama, Shinya Fujie, Tetsunori Kobayashi |
HRI | 3 |
| 2009 | Conversation robot participating in and activating a group communication
Shinya Fujie, Yoichi Matsuyama, Hikaru Taniyama, Tetsunori Kobayashi |
INTERSPEECH | 1 |
| 2009 | Upper-Body Contour Extraction Using Face and Body Shape Variance Information
Kazuki Hoshiai, Shinya Fujie, Tetsunori Kobayashi |
PSIVT | 2 |
| 2008 | An ASM fitting method based on machine learning that provides a robust parameter initialization for AAM fittingabstractDue to their use of information contained in texture, active appearance models (AAM) generally outperform active shape models (ASM) in terms of fitting accuracy. Although many extensions and improvements over the original AAM have been proposed, on of the main drawbacks of AAMs remains its dependence on good initial model parameters to achieve accurate fitting results. In this paper, we determine the initial model parameters for AAM fitting with ASM fitting, and use machine learning techniques to improve the scope and accuracy of ASM fitting. Combining the precision of AAM fitting with the large radius of convergence of learned ASM fitting improves the results by an order of magnitude, as our empirical evaluation on a database of publicly available benchmark images demonstrates. Matthias Wimmer, Shinya Fujie, Freek Stulp, Tetsunori Kobayashi, Bernd Radig |
FG | 2 |
| 2007 | Extensible speech recognition system using proxy-agentabstractThis paper presents an extension framework for a speech recognition system. This framework is designed to use “Proxy-Agent,” a software component located between applications, speech recognition engines, and input devices. By taking advantage of its structural characteristics, Proxy-Agent can provide supplementary services for speech recognition systems as well as user extensions. A monitoring capability, a feedback capability, and an extension capability are implemented and presented in this paper. For the first prototype, we developed a data collection application and an application control system using Proxy-Agent. Through these developments, we verified the effectiveness of the data collection capability of Proxy-Agent, and the framework extension capability. Teppei Nakano, Shinya Fujie, Tetsunori Kobayashi |
ASRU | 2 |
| 2006 | MONEA: Message-oriented Networked-robot ArchitectureabstractThis paper proposes message-oriented networked-robot architecture, or MONEA, as an efficient development platform architecture for multifunctional robots. In order to avoid problems occurred in multifunctional robot developments, we design the architecture to fulfill the following three features. Firstly, it embodies the meta-architecture for networked-robots. Secondly, it supports bazaar-style development model. Finally, it doesn't require heavy weight middleware. To realize them, we developed an information sharing framework named networked-whiteboard model along with message passing framework via P2P virtual network. A development methodology using interest-oriented module groups and software patterns is also presented as a means to reduce complexity risks. A middleware is developed as an implementation of this architecture, and we verify the availability and effectivity of our platform through the development of dialogue robot for exhibition Teppei Nakano, Shinya Fujie, Tetsunori Kobayashi |
ICRA | 2 |
| 2005 | Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system
Shinya Fujie, Kenta Fukushima, Tetsunori Kobayashi |
INTERSPEECH | 1 |
| 2004 | Prosody based attitude recognition with feature selection and its application to spoken dialog system as para-linguistic informationabstractIn this paper, prosody-based attitude recognition and its application to a spoken dialog system are proposed. Paralinguistic information plays a important role in the human communication. We aimed to recognize the user’s attitude by prosody, and apply it to a spoken dialog system as para-linguistic information. In order to find important features to recognize the attitude from automatically extracted features, we applied some feature selection methods. Experimental results show the stepwise method, a combination of the forward selection method and the backward selection method, achieved the best recognition rate. Finally, the dialog system using the recognition results as para-linguistic information is shown. Shinya Fujie, Tetsunori Kobayashi, Daizo Yagi, Hideaki Kikuchi |
INTERSPEECH | 1 |
| 2001 | Modeling of conversational strategy for the robot participating in the group conversationabstractThis paper describes a strategy for the conversation system to take part in human-to-human group conversation. One big characteristic of the group conversation system is that it can choose whether to observe or to take turn in the conversation. We implement the computational model combined with speech and gaze recognizers to keep the rules in turn taking, and define an interruption decision strategy based on an analysis of human needs. And finally, we realized a humanfriendly group conversation system by combining multimodal information processing/expression abilities of humanoid robot ROBITA. Yosuke Matsusaka, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 2 |