Yoshihiro Yamazaki

dblp:39/6550 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Distilling Knowledge of Bidirectional Language Model for Scene Text Recognition
abstract
This paper proposes a knowledge distillation method for an external bidirectional language model trained by masked language modeling to achieve high accuracy in scene text recognition. In Asian languages such as Japanese, it is necessary to perform text recognition in units of multiple words or sentences rather than individual words because words are not separated by spaces, and so high-level linguistic knowledge is needed to recognize text correctly. To enhance linguistic knowledge, several methods that use an external language model have been proposed, but these methods fail to consider future context well in performing text recognition because they revise the text candidates yielded by autoregressive text recognition models, which consider mainly past context. To overcome this deficiency, our key idea is to enhance a text recognition model by utilizing knowledge of an external bidirectional language model trained by masked language modeling, which reflects not only past but also future context. So as to actively consider future context in text recognition, our proposed method introduces a distillation loss term that makes the output probability of the text recognition model closer to that of the bidirectional language model. Experiments on Japanese scene text recognition demonstrate the effectiveness of the proposed method.
Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima, Ryo Masumura
ICIP2
2023 Open-Set Recognition for Facial-Expression Recognition
abstract
We address distinguishing whether an input is a facial image by learning only a facial-expression recognition (FER) dataset. To avoid misclassification in FER, it is necessary to distinguish whether the input is a facial image. Unfortunately, collecting exhaustive non-face images is costly. Therefore, distinguishing whether the input is a facial image by learning only an FER dataset is important. A representative method for this task is learning reconstruction of only facial images and determining high-error samples between input images and reconstructed images as non-face images. However, reconstruction is difficult on facial images because such images contain detailed features. Our key idea to tackle the task without reconstruction is assuming that facial images will match several emotions, whereas non-face images will not match any emotion. Therefore, we propose a method for training a discriminator that determines whether the inputs and emotions match using counter-factual pairs in an FER dataset. A metric for the task is then obtained by taking into account each emotion in the posterior probability that inputs and emotions match, estimated by the discriminator. Experiments on the RAF-DB dataset vs. the Stanford Dogs dataset and AffectNet datasets showed the effectiveness of our method.
Mihiro Uchida, Shota Orihashi, Akihiko Takashima, Yoshihiro Yamazaki, Ryo Masumura
ICIP4
2022 Fully Shareable Scene Text Recognition Modeling for Horizontal and Vertical Writing
abstract
This paper proposes a simple and efficient method of joint scene text recognition for both horizontal and vertical writing. Recently, end-to-end scene text recognition using the Transformer-based autoregressive encoder-decoder model offers high recognition accuracy. Research into this method has mainly focused on horizontally written text, but in several Asian countries, texts are also written vertically. To efficiently train a recognition model for jointly recognizing horizontal and vertical writing, several methods have been proposed that partially share model components between each writing direction. However, this approach lowers training efficiency because non-shareable components are trained only on just horizontal or vertical writing data. To increase training efficiency, our key idea is to consider writing direction in the continuous space obtained by a fully shareable model for horizontal and vertical writing. To this end, our proposed method gives the writing direction as an initial token to the autoregressive decoder while sharing all components for each writing direction. Furthermore, to incorporate common features between each writing into the model, the proposed method predicts character count before predicting the character string. Experiments on Japanese scene text recognition demonstrate the effectiveness of the proposed method.
Shota Orihashi, Yoshihiro Yamazaki, Mihiro Uchida, Akihiko Takashima, Ryo Masumura
ICIP2
2022 End-to-End Joint Modeling of Conversation History-Dependent and Independent ASR Systems with Multi-History Training
Ryo Masumura, Yoshihiro Yamazaki, Saki Mizuno, Naoki Makishima, Mana Ihori, Mihiro Uchida, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Shota Orihashi, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando
INTERSPEECH2
2022 Interactive Co-Learning with Cross-Modal Transformer for Audio-Visual Emotion Recognition
Akihiko Takashima, Ryo Masumura, Atsushi Ando, Yoshihiro Yamazaki, Mihiro Uchida, Shota Orihashi
INTERSPEECH4
2021 Hierarchical Knowledge Distillation for Dialogue Sequence Labeling
abstract
This paper presents a novel knowledge distillation method for dialogue sequence labeling. Dialogue sequence labeling is a supervised learning task that estimates labels for each utterance in the target dialogue document, and is useful for many applications such as dialogue act estimation. Accurate labeling is often realized by a hierarchically-structured large model consisting of utterance-level and dialogue-level networks that capture the contexts within an utterance and between utterances, respectively. However, due to its large model size, such a model cannot be deployed on resource-constrained devices. To overcome this difficulty, we focus on knowledge distillation which trains a small model by distilling the knowledge of a large and high performance teacher model. Our key idea is to distill the knowledge while keeping the complex contexts captured by the teacher model. To this end, the proposed method, hierarchical knowledge distillation, trains the small model by distilling not only the probability distribution of the label classification, but also the knowledge of utterance-level and dialogue-level contexts trained in the teacher model by training the model to mimic the teacher model's output in each level. Experiments on dialogue act estimation and call scene segmentation demonstrate the effectiveness of the proposed method.
Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura
ASRU2
2021 Neural Spoken-Response Generation Using Prosodic and Linguistic Context for Conversational Systems
Yoshihiro Yamazaki, Yuya Chiba, Takashi Nose, Akinori Ito
Interspeech1
2021 Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor Languages
abstract
This paper presents a novel training method for end-to-end scene text recognition. End-to-end scene text recognition offers high recognition accuracy, especially when using the encoder-decoder model based on Transformer. To train a highly accurate end-to-end model, we need to prepare a large image-to-text paired dataset for the target language. However, it is difficult to collect this data, especially for resource-poor languages. To overcome this difficulty, our proposed method utilizes well-prepared large datasets in resource-rich languages such as English, to train the resource-poor encoder-decoder model. Our key idea is to build a model in which the encoder reflects knowledge of multiple languages while the decoder specializes in knowledge of just the resource-poor language. To this end, the proposed method pre-trains the encoder by using a multilingual dataset that combines the resource-poor language’s dataset and the resource-rich language’s dataset to learn language-invariant knowledge for scene text recognition. The proposed method also pre-trains the decoder by using the resource-poor language’s dataset to make the decoder better suited to the resource-poor language. Experiments on Japanese scene text recognition using a small, publicly available dataset demonstrate the effectiveness of the proposed method.
Shota Orihashi, Yoshihiro Yamazaki, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, Ryo Masumura
MMAsia2
2020 Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness
abstract
There are high expectations for multimodal dialog systems that can make natural small talk with facial expressions, gestures, and gaze actions as next-generation dialog-based systems. Two important roles of the chat-talk system are keeping the user engaged and establishing rapport. Many studies have conducted user evaluations of such systems, some of which reported that considering the relationship with the user is an effective way to improve the subjective evaluation. To facilitate research of such dialog systems, we are currently constructing a large-scale multimodal dialog corpus focusing on the relationship between speakers. In this paper, we describe the data collection and annotation process, and analysis of the corpus collected in the early stage of the project. This corpus contains 19,303 utterances (10 hours) from 19 pairs of participants. A dialog act tag is annotated to each utterance by two annotators. We compare the frequency and the transition probability of the tags between different closeness levels to help construct a dialog system for establishing a relationship with the user.
Yoshihiro Yamazaki, Yuya Chiba, Takashi Nose, Akinori Ito
LREC1
2020 Physical Scattering Interpretation of POLSAR Coherency Matrix by Using Compound Scattering Phenomenon
abstract
The 2 × 2 relative scattering matrix [S] is characterized by five elements (three amplitudes and two relative phases). A more suitable representation of the data in terms of power expression for target identification exists in the 3 × 3 Paulibased polarimetric coherency matrix [T]. The aim of this work is to expand the physical interpretation of [T] elements in terms of a physical scattering mechanism, which has not yet been achieved completely. Compound scattering matrix nature is considered to interpret the elements of [T]. Compound scattering phenomenon, and thereby generated matrices, are verified with measured compound scattering matrices in an anechoic chamber by comparing measured and theoretical polarization signature responses. The behavior of measured polarization signature(s) is consistent with the theoretical/hypothetical polarization signature(s). It is also observed that the combination of two or more dipole-type targets formed the scattering generators and helped to reveal the unknown scattering mechanisms in the coherency matrix. Furthermore, the existence of compound scattering phenomenon is also justified with airborne and spaceborne fully polarimetric synthetic aperture radar (POLSAR) data by implementing recent physical model-based scattering power decomposition methods.
Gulab Singh, Shradha Mohanty, Yoshihiro Yamazaki, Yoshio Yamaguchi
IEEE Trans. Geosci. Remote. Sens.3