Satoshi Nakagawa

dblp:54/9296 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Retrieval-enhanced, Adaptively Collaborative, and Temporal-aware user behavior comprehension for LLM-based sequential recommendation
Zheng Hu 0001, Yongsen Pan, Zetao Li 0002, Satoshi Nakagawa, Jiawen Deng 0006, Shimin Cai, Fuji Ren
Inf. Process. Manag.5
2025 Bridging the User-side Knowledge Gap in Knowledge-aware Recommendations with Large Language Models
abstract
In recent years, knowledge graphs have been integrated into recommender systems as item-side auxiliary information, enhancing recommendation accuracy. However, constructing and integrating structural user-side knowledge remains a significant challenge due to the improper granularity and inherent scarcity of user-side features. Recent advancements in Large Language Models (LLMs) offer the potential to bridge this gap by leveraging their human behavior understanding and extensive real-world knowledge. Nevertheless, integrating LLM-generated information into recommender systems presents challenges, including the risk of noisy information and the need for additional knowledge transfer. In this paper, we propose an LLM-based user-side knowledge inference method alongside a carefully designed recommendation framework to address these challenges. Our approach employs LLMs to infer user interests based on historical behaviors, integrating this user-side information with item-side and collaborative data to construct a hybrid structure: the Collaborative Interest Knowledge Graph (CIKG). Furthermore, we propose a CIKG-based recommendation framework that includes a user interest reconstruction module and a cross-domain contrastive learning module to mitigate potential noise and facilitate knowledge transfer. We conduct extensive experiments on three real-world datasets to validate the effectiveness of our method. Our approach achieves state-of-the-art performance compared to competitive baselines, particularly for users with sparse interactions.
Zheng Hu 0001, Ziyun Jiao, Satoshi Nakagawa, Jiawen Deng 0006, Shimin Cai, Tao Zhou 0001, Fuji Ren
AAAI4
2025 Bridging Divides through Empathetic Robot Dialogue in Remote Interpersonal Conflict
abstract
In increasingly diverse and digitally connected societies, AI is being called upon not only to support communication, but to mediate it ethically—especially when conflicting values and psychological safety are at stake. While diversity is a cornerstone of well-being, it also presents challenges for mutual understanding—particularly in the absence of nonverbal cues and psychological safety. In this study, we propose a novel Human–Robot Interaction (HRI) framework wherein a social robot mediates dialogue between individuals by transforming conflict-prone utterances into empathic expressions in real time.Our system employs the BOCCO emo robot and a large language model (LLM) to reinterpret user speech, removing confrontational tones while preserving intent, and delivering the rephrased message to the interlocutor. Experimental results demonstrate that this empathic mediation significantly enhances partner impressions, perceived warmth, and openness, while reducing communication-related stress, indicating that empathic responses can initiate a reciprocal shift in dialogue tone.Beyond technical effectiveness, this approach raises important ethical implications: it suggests a new form of interaction design that respects individual expression while facilitating constructive dialogue. Rather than suppressing conflict, the robot reframes it in ways that support psychological safety and deepen engagement. This work contributes to a growing body of research that positions robots not merely as tools or agents, but as ethical mediators—capable of fostering mutual respect and enhancing well-being in human communication.
Satoshi Nakagawa, Misato Nihei
RO-MAN1
2025 Enhanced Emotion Recognition in Conversations Through Hybrid Context Encoding and Latent Dependency Mining
abstract
Emotion recognition in conversations (ERC) is a pivotal component of affective computing, involving a common two-stage paradigm where pre-trained language models first extract context-independent features, followed by the encoding of contextual information and the modeling of emotional dependencies. This paradigm faces two challenges: (1) Existing methods struggle to capture both the intra-dialogue emotional continuity and the inter-dialogue semantic similarity. (2) The complexity of emotional elicitation processes gives rise to entangled dependencies, termed “latent dependencies”, which are difficult for current methods to detect and analyze. To overcome these challenges, we propose a Hybrid-Context Encoder with an Automated Latent Dependency Mining model for ERC. Specifically, we examine the emotional continuity and the semantic similarity from the standpoint of context encoders. We experimentally find that context encoders with different architectures exhibit distinct benefits. Based on these findings, we design a hybrid contextual encoding module that effectively combines the strengths of various encoders. Additionally, we design a lightweight generative module for latent dependency mining that autonomously generates a context mask, enabling the effective discovery of latent dependencies. We conduct extensive experiments on three datasets in the text modality. Our model achieves the best performance, which validates the superiority of our approach.
Zheng Hu 0001, Jiawen Deng 0006, Satoshi Nakagawa, Yan Zhuang 0002, Shimin Cai, Fuji Ren
IEEE Trans. Affect. Comput.3
2025 Text-Guided Reconstruction Network for Sentiment Analysis With Uncertain Missing Modalities
abstract
Multimodal Sentiment Analysis (MSA) is an attractive research that aims to integrate sentiment expressed in textual, visual, and acoustic signals. There are two main problems in the existing methods: 1) the dominant role of the text is underutilization in unaligned multimodal data, and 2) the modality under uncertain missing feature is not sufficiently explored. This paper proposes a Text-guided Reconstruction Network (TgRN) for MSA with uncertain missing modalities in non-aligned sequences. The TgRN network includes three primary modules: Text-guided Extraction Module (TEM), Reconstruction Module (RM) and Text-guided Fusion Module (TFM). First, the TEM consists of the text-guided cross attention units and self-attention units to capture inter-modal features and intra-modal features, respectively. Second, leveraging enhanced attention units and a three-way squeeze-and-excitation block, the RM is designed to learn semantic information from incomplete data and reconstruct missing modality features. Third, the TFM utilizes a progressive modality-mixing adaptation gate to explore the dynamic correlations between nonverbal and verbal modalities, effectively addressing the modality gap issue. Finally, under the supervision of sentiment prediction loss and reconstruction loss, the TgRN effectively processes both uncertain missing-modality conditions and ideal complete modality conditions. Extensive experiments on CMU-MOSI and CH-SIMS demonstrate that our proposed method outperforms state-of-the-art approaches.
Piao Shi, Min Hu 0010, Satoshi Nakagawa, Xiangming Zheng, Xuefeng Shi, Fuji Ren
IEEE Trans. Affect. Comput.3
2025 Hierarchical Denoising for Robust Social Recommendation
abstract
Social recommendations leverage social networks to augment the performance of recommender systems. However, the critical task of denoising social information has not been thoroughly investigated in prior research. In this study, we introduce a hierarchical denoising robust social recommendation model to tackle noise at two levels: 1) intra-domain noise, resulting from user multi-faceted social trust relationships, and 2) inter-domain noise, stemming from the entanglement of the latent factors over heterogeneous relations (e.g., user-item interactions, user-user trust relationships). Specifically, our model advances a preference and social psychology-aware methodology for the fine-grained and multi-perspective estimation of tie strength within social networks. This serves as a precursor to an edge weight-guided edge pruning strategy that refines the model's diversity and robustness by dynamically filtering social ties. Additionally, we propose a user interest-aware cross-domain denoising gate, which not only filters noise during the knowledge transfer process but also captures the high-dimensional, nonlinear information prevalent in social domains. We conduct extensive experiments on three real-world datasets to validate the effectiveness of our proposed model against state-of-the-art baselines. We perform empirical studies on synthetic datasets to validate the strong robustness of our proposed model.
Zheng Hu 0001, Satoshi Nakagawa, Yan Zhuang 0002, Jiawen Deng 0006, Shimin Cai, Tao Zhou 0001, Fuji Ren
IEEE Trans. Knowl. Data Eng.2
2024 MSSTNet: A Multi-Scale Spatio-Temporal CNN-Transformer Network for Dynamic Facial Expression Recognition
abstract
Unlike typical video action recognition, Dynamic Facial Expression Recognition (DFER) does not involve distinct moving targets but relies on localized changes in facial muscles. Addressing this distinctive attribute, we propose a MultiScale Spatio-temporal CNN-Transformer network (MSST-Net). Our approach takes spatial features of different scales extracted by CNN and feeds them into a Multi-scale Embedding Layer (MELayer). The MELayer extracts multi-scale spatial information and encodes these features before sending them into a Temporal Transformer (T-Former). The T-Former simultaneously extracts temporal information while continually integrating multi-scale spatial information. This process culminates in the generation of multi-scale spatio-temporal features that are utilized for the final classification. Our method achieves state-of-the-art results on two in-the-wild datasets. Furthermore, a series of ablation experiments and visualizations provide further validation of our approach's proficiency in leveraging spatio-temporal information within DFER.
Linhuang Wang, Satoshi Nakagawa, Fuji Ren
ICASSP4
2024 Self Decoupling-Reconstruction Network for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) poses significant challenges due to various imaging conditions, including diverse head poses, lighting conditions, resolutions, and occlusions. Additionally, different personal attributes such as age, gender, and racial background further contribute to the complexity of FER. To accurately extract meaningful expression features amidst these interfering factors to enhance recognition accuracy and the model’s generalization, we propose a Self Decoupling-Reconstruction Network (SDRNet). Specifically, our approach involves two learning processes. In the first phase, the network is trained to decouple facial images with expressions into expression and neutral components. This process involves reconstructing neutral facial images and the original input, ensuring the preservation of meaningful expression components devoid of interference in the decoupling process. In the second learning phase, we employ simple convolutional neural networks (CNNs) to recognize the extracted expression components. Our method has achieved state-of-the-art results across multiple widely used datasets, providing substantial evidence of its effectiveness. Additionally, we demonstrate the robust generalization performance of our approach through cross-database evaluations.
Linhuang Wang, Hai-Tao Yu 0003, Yunong Wu, Kazuyuki Matsumoto, Satoshi Nakagawa, Fuji Ren
IJCNN7
2024 Emergence of Grounded Language Representations for Continuous Object Properties Through Decentralized Embodied Learning
Satoshi Nakagawa, Tomohiro Tanikawa, Yasuo Kuniyoshi
PRICAI (2)2
2024 Aspect based sentiment analysis with instruction tuning and external knowledge enhanced dependency graph
Xuefeng Shi, Fuji Ren, Piao Shi, Satoshi Nakagawa
Appl. Intell.5
2024 A multi-attention and depthwise separable convolution network for medical image segmentation
abstract
Automatic medical image segmentation method is highly needed to help experts in lesion segmentation. The deep learning technology emerging has profoundly driven the development of medical image segmentation. While U-Net and attention mechanisms are widely utilized in this field, the application of attention, albeit successful in natural scene image segmentation, tends to inflate the number of model parameters and neglects the potential for feature fusion between different convolutional layers. In response to these challenges, we present the Multi-Attention and Depthwise Separable Convolution U-Net (MDSU-Net), designed to enhance feature extraction. The multi-attention aspect of our framework integrates dual attention and attention gates, adeptly capturing rich contextual details and seamlessly fusing features across diverse convolutional layers. Additionally, our encoder integrates a depthwise separable convolution layer, streamlining the model’s complexity without sacrificing its efficacy, ensuring versatility across various segmentation tasks. The results demonstrate that our method outperforms state-of-the-art across three diverse medical image datasets.
Fuji Ren, Huimin Lu 0001, Satoshi Nakagawa, Xiao Shan
Neurocomputing5
2024 Enhancing cross-market recommendations by addressing negative transfer and leveraging item co-occurrences
Zheng Hu 0001, Satoshi Nakagawa, Shimin Cai, Fuji Ren, Jiawen Deng 0006
Inf. Syst.2
2023 Celebrity-aware Graph Contrastive Learning Framework for Social Recommendation
abstract
Social networks exhibit a distinct "celebrity effect" whereby influential individuals have a more significant impact on others compared to ordinary individuals, unlike other network structures such as citation networks and knowledge graphs. Despite its common occurrence in social networks, the celebrity effect is frequently overlooked by existing social recommendation methods when modeling social relationships, thereby hindering the full exploitation of social networks to mine similarities between users. In this paper, we fill this gap and propose a Celebrity-aware Graph Contrastive Learning Framework for Social Recommendation (CGCL), which explicitly models the celebrity effect in the social domain. Technically, we measure the different influences of celebrity and ordinary nodes by mining social network structure features, such as closeness centrality. To model the celebrity effect in social networks, we design a novel user-user impact-aware aggregation method, which incorporates the celebrity-aware influence information into the message propagation process. Additionally, we design a graph neural network-based framework which incorporates social semantics into the user-item interaction modeling with contrastive learning-enhanced data augmentation. The experimental results on three real-world datasets show the effectiveness of the proposed framework. We conduct ablation experiments to prove that the key components of our model benefit the recommendation performance improvement.
Zheng Hu 0001, Satoshi Nakagawa, Yu Gu 0003, Fuji Ren
CIKM2
2023 CenterMatch: A Center Matching Method for Semi-supervised Facial Expression Recognition
Linhuang Wang, Satoshi Nakagawa, Fuji Ren
PRCV (6)3
2023 Impact of QOL-Based Robot Counseling on Older Adults' QOL Improvement
abstract
This study aimed to develop a counseling robot to improve the quality of life (QOL) of older adults and verify its effectiveness. QOL is a comprehensive indicator that includes physical, mental, and social aspects, and gerontechnology aims to improve the autonomy and QOL of older adults. In recent years, utilizing robots to address the shortage of caregivers has been increasingly studied. We developed a counseling robot that estimates the QOL of older adults in real time, and generates appropriate and empathetic responses according to the estimated QOL. The results obtained from a one-week interaction experiment with a counseling robot targeted toward older adults indicated a significant improvement in the mental aspect of QOL, and the use of cognitive behavioral therapy and empathetic responses was inferred to facilitate self-disclosure and enhance the effectiveness of counseling. Additionally, advice based on the QOL estimation results contributed to organizing the thoughts of older adults and further improved their mental health. This study not only contributes to improving the QOL of older adults but also suggests that robots that understand and appropriately respond to individuals can facilitate continuous relationship building. This study may also serve as a guide for promoting the introduction of information and communication technology into welfare facilities.
Satoshi Nakagawa, Kana Naruse, Ryoga Endo, Yasuo Kuniyoshi
SMC1
2023 DEU-Net: Dual Encoder U-Net for 3D Medical Image Segmentation
abstract
As medical image analysis equipment has evolved and gained popularity, MRI has taken the forefront in radiological imaging. Since 3D medical images are more complex and contextual features are more difficult to capture, Transformer needs to be introduced to enhance the feature extraction capabilities of the network. We propose Dual Encoder U-Net (DEU-Net), which uses Transformer and CNN to extract medical image features in the encoder. Transformer is a pre-trained model in BTCV, which improves its ability to capture contextual features of medical images and increases the learning speed. To fuse the two kinds of features, we propose a Dual Feature Fusion Module (DFFM) to fuse the features extracted from the Transformer and CNN respectively, making full use of the feature extraction capabilities of the two extractors for 3D medical image. The results demonstrate that DEU-Net outperforms state-of-the-art for three segmentation tasks on the BraTS 2020 dataset.
Fuji Ren, Satoshi Nakagawa, Xiao Shan
TrustCom4
2020 Estimation of Mental Health Quality of Life using Visual Information during Interaction with a Communication Agent
abstract
It is essential for a monitoring system or a communication robot that interacts with an elderly person to accurately understand the user's state and generate actions based on their condition. To ensure elderly welfare, quality of life (QOL) is a useful indicator for determining human physical suffering and mental and social activities in a comprehensive manner. In this study, we hypothesize that visual information is useful for extracting high-dimensional information on QOL from data collected by an agent while interacting with a person. We propose a QOL estimation method to integrate facial expressions, head fluctuations, and eye movements that can be extracted as visual information during the interaction with the communication agent. Our goal is to implement a multiple feature vectors learning estimator that incorporates convolutional 3D to learn spatiotemporal features. However, there is no database required for QOL estimation. Therefore, we implement a free communication agent and construct our database based on information collected through interpersonal experiments using the agent. To verify the proposed method, we focus on the estimation of the "mental health" QOL scale, which is the most difficult to estimate among the eight scales that compose QOL based on a previous study. We compare the four estimation accuracies: single-modal learning using each of the three features, i.e., facial expressions, head fluctuations, and eye movements and multiple feature vectors learning integrating all the three features. The experimental results show that multiple feature vectors learning has fewer estimation errors than all the other single-modal learning, which uses each feature separately. The experimental results for evaluating the difference between the estimated QOL score by the proposed method and the actual QOL score calculated by the conventional method also show that the average error is less than 10 points and, thus, the proposed system can estimate the QOL score. Thus, it is clear that the proposed new approach for estimating human conditions can improve the quality of human-robot interactions and personalized monitoring.
Satoshi Nakagawa, Shogo Yonekura, Hoshinori Kanazawa, Satoshi Nishikawa, Yasuo Kuniyoshi
RO-MAN1
2020 New telecare approach based on 3D convolutional neural network for estimating quality of life
abstract
Quality of life (QoL) is an effective index of well-being, including physical health, aspect of social activity, and mental state of individuals. A new approach that uses a deep-learning architecture to estimate the score of a user's QoL is presented. This system was built using a combination of a 3D convolutional neural network and a support vector machine for multimodal data. In order to evaluate the accuracy of the estimation system, three experiments were conducted. Before these experiments, ten hours of audio and video data were collected from healthy participants during a natural-language conversation with a conversational agent we implemented. In the first experiment, the QoL question-answer estimation experiment, the accuracy of “Physical functioning,” which is one of the eight scales that constitute QoL, reached 84.0%. In the second experiment, the QoL-score-regression experiment, in which the scores of each scale were directly estimated, the distribution of the difference between the actual score and the estimated results, known as error, was investigated. These results imply that the features necessary for QoL estimation can be extracted from audio and video data, except for the “Mental Health” domain. One of the reasons why it was difficult to estimate the “Mental Health” scale may be that the learning framework could not extract an appropriate feature for estimation. Therefore, we estimated “Mental Health” by focusing on eye movement. From the result, it was proven that estimation is possible, and the proposed system using multimodal data demonstrated its effectiveness for estimation for all eight scales that constitute QoL and for extracting high-dimensional information regarding the QoL of a human, including their satisfaction level towards daily life and social activities. Finally, suggestions and discussions regarding the plausible behavior of the estimation results were made from the viewpoint of human–agent interaction in the field of elderly welfare.
Satoshi Nakagawa, Daiki Enomoto, Shogo Yonekura, Hoshinori Kanazawa, Yasuo Kuniyoshi
Neurocomputing1
2016 Enhanced decomposition-based many-objective optimization using supplemental weight vectors
abstract
In evolutionary multi-objective optimization, each solution in the population generally has two roles. The first one is to approximate a part of the Pareto front, and the second one is to be a variable information resource to generate offspring. In many-objective optimization involving four or more conflicting objectives, solutions in the population have to be sparsely distributed in the objective space and the variable space to approximate a high-dimensional Pareto front, and each solution faces the difficulty to play the second role since variables are drastically individualized in the population. To overcome this problem, we focus on MOEA/D algorithm framework and propose a method to introduce supplemental weight vectors and solutions which maintain variable information resource to enhance the solution search for each part of the Pareto front. Experimental results using many-objective knapsack problems show that the supplemental weight vectors and solutions improves the search performance of MOEA/D by improving the diversity of the obtained solutions.
Hiroyuki Sato 0003, Satoshi Nakagawa, Minami Miyakawa, Keiki Takadama
CEC2
2010 A real-time system of distributed video coding
abstract
This paper presents a real-time system of distributed video coding (DVC). DVC is a current video compression paradigm. The decoding process of DVC is normally complex, which causes difficulty in real-time implementation. To address this problem, we propose a new configuration of DVC with three methods: simple rate control without the feedback channel, simple transmitting of dynamic range and simple bidirectional motion estimation to reduce complexity. Then we implement the system with parallelization techniques. We also develop the encoder for a low power processor. Experimental results show that the encoder on i.MX31 400 MHz could operates at about CIF 13 fps, and the decoder on Core 2 Quad 2.83 GHz operates at more than CIF 30 fps.
Kazuhito Sakomizu, Takahiro Yamasaki, Satoshi Nakagawa, Takashi Nishi
PCS3
2002 Automatic user-adaptive speaking rate selection for information delivery
abstract
Today there are many services which provide information over the phone using a prerecorded or synthesized voice. These voices are invariant in speed. Humans giving information over the telephone, however, tend to adapt the speed of their presentation to suit the needs of the listener. This paper presents a preliminary model of this adaptation. In a corpus of simulated directory assistance dialogs the operator 's speed in number-giving correlates with the speed of the user's initial response and with the user's speaking rate. Multiple regression gives a formula which predicts appropriate speaking rates, and these predictions correlate (.46) with the speeds observed in good dialogs in the corpus. An experiment with 18 subjects suggests that users prefer a system which adapts its speed to the user in this way.
Nigel Ward, Satoshi Nakagawa
INTERSPEECH2
1996 Automated detection of human for visual surveillance system
abstract
This paper describes a robust and reliable method of human detection for visual surveillance systems. The merit of this method is to use simple shape parameters of silhouette patterns to classify humans from other moving objects such as butterflies and autonomous factory vehicles. An extra function based on the brightness level transformation is used to extract the precise shape of the silhouette patterns. An approach to overcome the occlusions of humans is also proposed. We tested our method for 2,500 images (1,100 from humans and 1,400 from other moving objects). Our test system detected the humans at the rate of 98% (=1,077/1,100) and judged 92% (=1,283/1,400) of the other moving objects as non-humans.
Yuji Kuno, Nakayama Watanabe, Yoshinori Shimosakoda, Satoshi Nakagawa
ICPR4