Gen Hattori

dblp:61/872 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
8since 2021 · last 2023
0000-0001-5636-1065ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Deep Learning Pipeline for Spotting Macro- and Micro-expressions in Long Video Sequences Based on Action Units and Optical Flow
Bo Yang 0057, Kazushi Ikeda, Gen Hattori, Masaru Sugano, Yusuke Iwasawa, Yutaka Matsuo
Pattern Recognit. Lett.4
2023 Learning Explicit and Implicit Dual Common Subspaces for Audio-visual Cross-modal Retrieval
abstract
Audio-visual tracks in video contain rich semantic information with potential in many applications and research. Since the audio-visual data have inconsistent distributions and because of the heterogeneous nature of representations, the heterogeneous gap between modalities makes them impossible to compare directly. To bridge the modality gap, a frequently adopted approach is to simultaneously project audio-visual data into a common subspace to capture the commonalities and characteristics of modalities for measurement, which has been extensively studied in relation to the issues of modality-common and modality-specific feature learning in previous research. However, it is difficult for existing methods to address the tradeoff between both issues; e.g., the modality-common feature is learned from the latent commonalities of audio-visual data or the correlated features as aligned projections, in which the modality-specific feature can be lost. To solve the tradeoff, we propose a novel end-to-end architecture, which synchronously projects audio-visual data into the explicit and the implicit dual common subspaces. The explicit subspace is used to learn modality-common features and reduce the modality gap of explicitly paired audio-visual data, where the representation-specific details are abandoned to retain the common underlying structure of audio-visual data. The implicit subspace is used to learn modality-specific features, where each modality privately pulls apart the feature distances between different categories to maintain the category-based distinctions, by minimizing the distance between audio-visual features and corresponding labels. The comprehensive experimental results on two audio-visual datasets, VEGAS and AVE, demonstrate that our proposed model for using two different common subspaces for audio-visual cross-modal learning is effective and significantly outperforms the state-of-the-art cross-modal models that learn features from a single common subspace by 4.30% and 2.30% in terms of average MAP on the VEGAS and AVE datasets, respectively.
Donghuo Zeng, Gen Hattori, Yi Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Face-mask-aware Facial Expression Recognition based on Face Parsing and Vision Transformer
Bo Yang 0057, Kazushi Ikeda, Gen Hattori, Masaru Sugano, Yusuke Iwasawa, Yutaka Matsuo
Pattern Recognit. Lett.4
2021 Face Mask Aware Robust Facial Expression Recognition During The Covid-19 Pandemic
abstract
Wearing face masks is considered an effective means of preventing the transmission of coronavirus during the COVID-19 pandemic. Facial expression recognition (FER) under partial occlusion, especially with face masks, makes it a challenging task in the research area of computer vision. In this paper, we propose a two-stage attention model to improve the accuracy of face-mask-aware FER: In stage 1, we train the masked/unmasked binary deep classifier, which can generate attention heatmaps to roughly distinguish the masked facial parts from the unobstructed region. In stage 2, we train the FER classifier, which is guided to pay more attention to the region that is essential to the facial expression classification, and both occluded and non-occluded regions are taken into consideration but reweighed. The proposed method outperforms other state-of-the-art occlusion-aware FER methods on face-mask-aware FER datasets, whether in the wild or in the laboratory.
Gen Hattori
ICIP3
2021 SHECS: A Local Smart Hands-free Elderly Care Support System on Smart AR Glasses with AI Technology
abstract
Some elderly care homes attempt to remedy the shortage of skilled caregivers and provide long-term care for the elderly residents, by enhancing the management of the care support system with the aid of smart devices such as mobile phones and tablets. Since mobile phones and tablets lack the flexibility required for laborious elderly care work, smart AR glasses have already been considered. Although lightweight smart AR devices with a transparent display are more convenient and responsive in an elderly care workplace, fetching data from the server through the Internet results in network congestion not to mention the limited display area. To devise portable smart AR devices that operate smoothly, we first present a no-keepalive-Internet and privacy-compliant required smart hands-free elderly care support system that employs smart glasses with facial recognition and text-to-speech synthesis technologies. Our support system utilizes automatic lightweight facial recognition to identify residents, and information about each resident in question can be obtained hands-free link with a local database. Moreover, a resident information can be displayed immediately on just a portion of the AR smart glasses on the spot. Due to the limited size of the display area, it cannot show all the necessary information. We therefore exploit synthesized voice in the system to read out the elderly care-related information in its entirety. By using the support system, caregivers can gain an understanding of each resident condition immediately, instead of having to devote considerable time in advance in obtaining the complete information of all elderly residents, especially of new residents. Our experiments on this support system were conducted at an elderly care home in Tokyo. Our lightweight facial recognition model achieved high accuracy with fewer model parameters than current state-of-the-art methods. The validation rate of our facial recognition system in the practical use scenario was 99.3% or higher with the false accept rate of 0.001, and caregivers rated the acceptability at 3.6 (5 levels) or higher.
Donghuo Zeng, Tomohiro Obara, Akeri Okawa, Nobuko Iino, Gen Hattori, Ryoichi Kawada, Yasuhiro Takishima
ISM7
2021 Sync Glass: Virtual Pouring and Toasting Experience with Multimodal Presentation
abstract
One of the challenges of non-face-to-face communication is the absence of the haptic dimension. To solve this, a haptic communication system via the Internet has been proposed. The system has to be designed in such a way that it does not create discomfort during general use. The "Sync Glass" that we have developed transmits and presents the feeling of pouring a drink and making a toast accompanied by haptic, sound and visual effects. The device is designed to resemble a glass cup and, moreover, each action, including drinking and making a toast is performed in the customary way, making its use more acceptable to users. In the internal user demonstrations we performed, the experience has been reviewed with participants saying that "the feeling of pouring is so realistic", "so enjoyable!", and similar affirmative statements.
Yuki Tajima, Toshiharu Horiuchi, Gen Hattori
ACM Multimedia3
2021 Facial Action Unit-based Deep Learning Framework for Spotting Macro- and Micro-expressions in Long Video Sequences
abstract
In this paper, we utilize facial action units (AUs) detection to construct an end-to-end deep learning framework for the macro- and micro-expressions spotting task in long video sequences. The proposed framework focuses on individual components of facial muscle movement rather than processing the whole image, which eliminates the influence of image change caused by noises, such as body or head movement. Compared with existing models deploying deep learning methods with classical Convolutional Neural Network (CNN) models, the proposed framework utilizes Gated Recurrent Unit (GRU) or Long Short-term Memory (LSTM) or our proposed Concat-CNN models to learn the characteristic correlation between AUs of distinctive frames. The Concat-CNN uses three convolutional kernels with different sizes to observe features of different duration and emphasizes both local and global mutation features by changing dimensionality (max-pooling size) of the output space. Our proposal achieves state-of-the-art performance from the aspect of overall F1-scores: 0.2019 on CAS(ME)2-cropped, 0.2736 on SAMM Long Video, and 0.2118 on CAS(ME)2, which not only outperforms the baseline but is also ranked the 3rd of FME challenge 2021 for combined datasets of CAS(ME)2-cropped and SAMM-LV.
Zhiguang Zhou, Megumi Komiya, Koki Kishimoto, Keisuke Nonaka, Toshiharu Horiuchi, Satoshi Komorita, Gen Hattori, Sei Naito, Yasuhiro Takishima
ACM Multimedia10
2021 TV-watching Companion Robot Supported by Open-domain Chatbot "KACTUS"
abstract
Watching TV once encouraged generations of families and friends [11] to communicate and share empathy. However, the Internet is changing how we watch TV and reducing interaction, leading to problems such as lack of self-control and inadequate communication skills [17]. To understand the conversations while watching TV, we design a scheme based on human conversational behavior [2], and then develop a prototype of TV-watching companion robot supported by the chatbot “KACTUS” [20]. The robot generates a disclosure utterance (e.g., ”I like elephants”) with extracted keywords from the TV program in “TV-watching mode” and uses a cross-topic dialogue management method from “KACTUS” with question utterance to respond with rich conversations in ”Conversation mode”. The robot switches between these two modes at a preset ratio (TV-watching:3, Conversation:1) and behaves like a human enjoying TV-watching. The result of initial experiment shows that three groups of participants enjoyed talking with the robot and the question about their interests in the robot were rated 6.5 (7-levels: ascending from ”extremely disagree” to ”extremely agree”).
Donghuo Zeng, Gen Hattori, Yasuhiro Takishima, Yuta Hagio, Marina Kamimura, Yuta Hoshi, Yutaka Kaneko, Yusei Nishimoto
MUM4
2020 LDNN: Linguistic Knowledge Injectable Deep Neural Network for Group Cohesiveness Understanding
abstract
Group cohesiveness reflects the level of intimacy that people feel with each other, and the development of a dialogue robot that can understand group cohesiveness will lead to the promotion of human communication. However, group cohesiveness is a complex concept that is difficult to predict based only on image pixels. Inspired by the fact that humans intuitively associate linguistic knowledge accumulated in the brain with the visual images they see, we propose a linguistic knowledge injectable deep neural network (LDNN) that builds a visual model (visual LDNN) for predicting group cohesiveness that can automatically associate the linguistic knowledge hidden behind images. LDNN consists of a visual encoder and a language encoder, and applies domain adaptation and linguistic knowledge transition mechanisms to transform linguistic knowledge from a language model to the visual LDNN. We train LDNN by adding descriptions to the training and validation sets of the Group AFfect Dataset 3.0 (GAF 3.0), and test the visual LDNN without any description. Comparing visual LDNN with various fine-tuned DNN models and three state-of-the-art models in the test set, the results demonstrate that the visual LDNN not only improves the performance of the fine-tuned DNN model leading to an MSE very similar to the state-of-the-art model, but is also a practical and efficient method that requires relatively little preprocessing. Furthermore, ablation studies confirm that LDNN is an effective method to inject linguistic knowledge into visual models.
Yanan Wang 0002, Jinfa Huang, Gen Hattori, Yasuhiro Takishima, Shinya Wada, Rui Kimura, Satoshi Kurihara
ICMI4
2020 Advanced Multi-Instance Learning Method with Multi-features Engineering and Conservative Optimization for Engagement Intensity Prediction
abstract
This paper proposes an advanced multi-instance learning method with multi-features engineering and conservative optimization for engagement intensity prediction. It was applied to the EmotiW Challenge 2020 and the results demonstrated the proposed method's good performance. The task is to predict the engagement level when a subject-student is watching an educational video under a range of conditions and in various environments. As engagement intensity has a strong correlation with facial movements, upper-body posture movements and overall environmental movements in a given time interval, we extract and incorporate these motion features into a deep regression model consisting of layers with a combination of long short-term memory(LSTM), gated recurrent unit (GRU) and a fully connected layer. In order to precisely and robustly predict the engagement level in a long video with various situations such as darkness and complex backgrounds, a multi-features engineering function is used to extract synchronized multi-model features in a given period of time by considering both short-term and long-term dependencies. Based on these well-processed engineered multi-features, in the 1st training stage, we train and generate the best models covering all the model configurations to maximize validation accuracy. Furthermore, in the 2nd training stage, to avoid the overfitting problem attributable to the extremely small engagement dataset, we conduct conservative optimization by applying a single Bi-LSTM layer with only 16 units to minimize the overfitting, and split the engagement dataset (train + validation) with 5-fold cross validation (stratified k-fold) to train a conservative model. The proposed method, by using decision-level ensemble for the two training stages' models, finally win the second place in the challenge (MSE: 0.061110 on the testing set).
Yanan Wang 0002, Gen Hattori
ICMI4
2020 Facial Expression Recognition with the advent of face masks
abstract
With the worldwide spread of COVID-19, wearing face masks while interaction in public is becoming a common behavior to protect against infection. Thus, how to improve effectiveness of existing facial expression recognition (FER) technology on masked faces has become an urgent issue. However, there are no publicly available masked facial expression recognition datasets that take facial orientation into consideration. To address this issue, we propose a method that can add face masks to existing FER datasets automatically using differently shaped masks according to facial orientations. The FER models based on VGG19 and MobileNet are trained on public and private FER datasets added with mask. As part of our contribution, we collected real-world masked faces from the Internet using emotional keywords and constructed a masked FER test dataset for a fair performance evaluation. The experimental results show that training an FER model based on a simulated masked FER dataset is feasible.
Gen Hattori
MUM3
2019 Development and Evaluation of Japanese Text-to-speech Middleware for 32-Bit Microcontrollers
abstract
Japanese text-to-speech (TTS) middleware for 32-bit microcontrollers (MCUs) such as Arm Cortex-M4 has been developed. Our TTS middleware is based on HMM-based speech synthesis techniques and includes an analyzer to generate pronunciation from texts that consist of Kanji (ideographic) and Kana (syllabary) characters. The middleware has been highly optimized for MCUs with the succinct data structure for data compression, fixed-point arithmetic for fast processing and pipelined processing to reduce both the required RAM size and response time. In this study, it is demonstrated that a real-time TTS system implemented on a 14-pin DIP-size MCU board that consist mainly of an MCU and external serial NOR flash can synthesize 32 kHz-sampled speech sounds with quality comparable to that of the conventional implementation of the HMM-based speech synthesis. The peak current of the MCU board at that condition is approximately 15 mA.
Nobuyuki Nishizawa, Tomohiro Obara, Gen Hattori
ICASSP3
2018 Automatic Method to Build a Dictionary for Class-Based Translation Systems
Kohichi Takai, Gen Hattori, Keiji Yasuda, Panikos Heracleous, Akio Ishikawa, Kazunori Matsumoto, Fumiaki Sugaya
CICLing (1)2
2014 Feature Based Sentiment Analysis of Tweets in Multiple Languages
Maike Erdmann, Kazushi Ikeda, Hiromi Ishizaki, Gen Hattori, Yasuhiro Takishima
WISE (2)4
2013 Twitter user profiling based on text and community mining for market analysis
Kazushi Ikeda, Gen Hattori, Chihiro Ono, Hideki Asoh, Teruo Higashino
Knowl. Based Syst.2
2013 A content search system considering the activity and context of a mobile user
Mayu Iwata, Hiroki Miyamoto, Takahiro Hara, Daijiro Komaki, Kentaro Shimatani, Tomohiro Mashita, Kiyoshi Kiyokawa, Toshiaki Uemukai, Gen Hattori, Shojiro Nishio, Haruo Takemura
Pers. Ubiquitous Comput.9
2012 Hierarchical Training of Multiple SVMs for Personalized Web Filtering
Maike Erdmann, Duc Dung Nguyen, Tomoya Takeyoshi, Gen Hattori, Kazunori Matsumoto, Chihiro Ono
PRICAI4
2007 Robust web page segmentation for mobile terminal using content-distances and page layout information
abstract
The demand of browsing information from general Web pages using a mobile phone is increasing. However, since the majority of Web pages on the Internet are optimized for browsing from PCs, it is difficult for mobile phone users to obtain sufficient information from the Web. Therefore, a method to reconstruct PC-optimized Web pages for mobile phone users is essential. An example approach is to segment the Web page based on its structure, and utilize the hierarchy of the content element to regenerate a page suitable for mobile phone browsing. In our previous work, we have examined a robust automatic Web page segmentation scheme which uses the distance between content elements based on the relative HTML tag hierarchy, i.e., the number and depth of HTML tags in Web pages. However, this scheme has a problem that the content-distance based on the order of HTML tags does not always correspond to the intuitional distance between content elements on the actual layout of a Web page. In this paper, we propose a hybrid segmentation method which segments Web pages based on both the content-distance calculated by the previous scheme, and a novel approach which utilizes Web page layout information. Experiments conducted to evaluate the accuracy of Web page segmentation results prove that the proposed method can segment Web pages more accurately than conventional methods. Furthermore, implementation and evaluation of our system on the mobile phone prove that our method can realize superior usability compared to commercial Web browsers.
Gen Hattori, Keiichiro Hoashi, Kazunori Matsumoto, Fumiaki Sugaya
WWW1
2003 Making Java-Enabled Mobile Phone as Ubiquitous Terminal by Lightweight FIPA Compliant Agent Platform
abstract
We discuss the design issues on lightweight and FIPA compliant agent platform for Java-enabled mobile phones and describe the design of such agent platform. This platform changes Java-enabled mobile phones to ubiquitous terminals by providing place for agent applications. Combined with location services, it can be used for various ubiquitous services. We also show the performance comparison of the prototype with LEAP, another lightweight agent platform.
Gen Hattori, Satoshi Nishiyama, Chihiro Ono, Hiroki Horiuchi
PerCom1