Donghuo Zeng

dblp:193/8040 · DBLP profile ↗
← Back
17ranked-venue papers
13as first author
13since 2021 · last 2026
0000-0002-6425-6270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
Donghuo Zeng, Hao Niu 0001, Yanan Wang 0002, Masato Taya
ECIR (2)1
2026 Personality-Aware Reinforcement Learning for Persuasive Dialogue with LLM-Driven Simulation
Donghuo Zeng, Roberto Legaspi, Kazushi Ikeda
PERSUASIVE1
2025 Learning Hidden Causal Factors from Psychometrics Data Using Distributional Information
Roberto Legaspi, Xinshuai Dong, Donghuo Zeng, Yuewen Sun, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
CogSci3
2025 Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning
abstract
Metric learning projects samples into an embedded space, where similarities and dissimilarities are quantified based on their learned representations. However, existing methods often rely on label-guided representation learning, where representations of different modalities, such as audio and visual data, are aligned based on annotated labels. This approach tends to underutilize latent complex features and potential relationships inherent in the distributions of audio and visual data that are not directly tied to the labels, resulting in suboptimal performance in audio-visual embedding learning. To address this issue, we propose a novel architecture that integrates cross-modal triplet loss with progressive self-distillation. Our method enhances representation learning by leveraging inherent distributions and dynamically refining soft audio-visual alignments—probabilistic aligns between audio and visual data that capture the inherent relationships beyond explicit labels. Specifically, the model distills audio-visual distribution-based knowledge from annotated labels in a subset of each batch. This self-distilled knowledge is used to automatically generate soft-alignment labels for the remaining audio-visual samples. These soft-alignment labels are used to construct soft cross-modal triplets, which in turn are employed to fine-tune the model’s parameters. Experimental results on two audio-visual benchmark datasets demonstrate the effectiveness of our proposed method in the cross-modal retrieval task, achieving state-of-the-art performance with improvements of 2.13% and 1.82% on the AVE and VEGAS datasets, respectively, in terms of average MAP metrics.
Donghuo Zeng, Kazushi Ikeda
ICASSP1
2025 Generative Framework for Personalized Persuasion: Inferring Causal, Counterfactual, and Latent Knowledge
abstract
We hypothesize that optimal system responses emerge from adaptive strategies grounded in causal and counterfactual knowledge.Counterfactual inference allows us to create hypothetical scenarios to examine the effects of alternative system responses.We enhance this process through causal discovery, which identifies the strategies informed by the underlying causal structure that govern system behaviors.Moreover, we consider the psychological constructs and unobservable noises that might be influencing user-system interactions as latent factors.We show that these factors can be effectively estimated.We employ causal discovery to identify strategy-level causal relationships among user and system utterances, guiding the generation of personalized counterfactual dialogues.We model the user utterance strategies as causal factors, enabling system strategies to be treated as counterfactual actions.Furthermore, we optimize policies for selecting system responses based on counterfactual data.Our results using a real-world dataset on social good demonstrate significant improvements in persuasive system outcomes, with increased cumulative rewards validating the efficacy of causal discovery in guiding personalized counterfactual inference and optimizing dialogue policies for a persuasive dialogue system.
Donghuo Zeng, Roberto Legaspi, Yuewen Sun, Xinshuai Dong, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
UMAP1
2024 Anchor-aware Deep Metric Learning for Audio-visual Retrieval
abstract
Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performance of tasks like audio-visual cross-modal retrieval (AV-CMR). Recent works employ sampling methods to select impactful data points from the embedding space during training. However, the model training fails to fully explore the space due to the scarcity of training data points, resulting in an incomplete representation of the overall positive and negative distributions. In this paper, we propose an innovative Anchor-aware Deep Metric Learning (AADML) method to address this challenge by uncovering the underlying correlations among existing data points, which enhances the quality of the shared embedding space. Specifically, our method establishes a correlation graph-based manifold structure by considering the dependencies between each sample as the anchor and its semantically similar samples. Through dynamic weighting of the correlations within this underlying manifold structure using an attention-driven mechanism, Anchor Awareness (AA) scores are obtained for each anchor. These AA scores serve as data proxies to compute relative distances in metric learning approaches. Extensive experiments conducted on two audio-visual benchmark datasets demonstrate the effectiveness of our proposed AADML method, significantly surpassing state-of-the-art models. Furthermore, we investigate the integration of AA proxies with various metric learning methods, further highlighting the efficacy of our approach.
Donghuo Zeng, Yanan Wang 0002, Kazushi Ikeda, Yi Yu 0001
ICMR1
2024 Identifying Latent State-Transition Processes for Individualized Reinforcement Learning
abstract
The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions in healthcare to learning progress in education. As a result, different individuals may exhibit different state-transition processes. Understanding individualized state-transition processes is essential for optimizing individualized policies. In practice, however, identifying these state-transition processes is challenging, as individual-specific factors often remain latent. In this paper, we establish the identifiability of these latent factors and introduce a practical method that effectively learns these processes from observed state-action trajectories. Experiments on various datasets show that the proposed method can effectively identify latent state-transition processes and facilitate the learning of individualized RL policies.
Yuewen Sun, Biwei Huang, Yu Yao 0005, Donghuo Zeng, Xinshuai Dong, Songyao Jin, Roberto Legaspi, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
NeurIPS4
2024 Counterfactual Reasoning Using Predicted Latent Personality Dimensions for Optimizing Persuasion Outcome
Donghuo Zeng, Roberto Legaspi, Yuewen Sun, Xinshuai Dong, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
PERSUASIVE1
2023 Triplet Loss with Curriculum Learning for Audio-Visual Retrieval
abstract
The cross-modal retrieval models leverage the potential of triple loss optimization to learn robust embedding spaces. However, existing methods often train these models in a singular pass, overlooking the distinction between semi-hard and hard triples in the optimization process, which will lead to suboptimal model performance. In this paper, we introduce a novel approach rooted in curriculum learning to address this problem. We propose a two-stage training paradigm that guides the model’s learning process from semi-hard to hard triplets. In the first stage, the model is trained with a set of semi-hard triplets, starting from a low-loss base. Subsequently, in the second stage, the model mines the hardest triplet with the primary aim of mitigating the risk of overfitting by addressing the highest loss. Extensive experimental results conducted on the audio-visual dataset show a significant improvement of approximately 9.8% in terms of average Mean Average Precision (MAP) over the current state-of-the-art method, MSNSCA, for the Audio-Visual Cross-Modal Retrieval (AV-CMR) task on the AVE dataset, indicating the effectiveness of our proposed method.
Donghuo Zeng, Kazushi Ikeda
ISM1
2023 Learning Explicit and Implicit Dual Common Subspaces for Audio-visual Cross-modal Retrieval
abstract
Audio-visual tracks in video contain rich semantic information with potential in many applications and research. Since the audio-visual data have inconsistent distributions and because of the heterogeneous nature of representations, the heterogeneous gap between modalities makes them impossible to compare directly. To bridge the modality gap, a frequently adopted approach is to simultaneously project audio-visual data into a common subspace to capture the commonalities and characteristics of modalities for measurement, which has been extensively studied in relation to the issues of modality-common and modality-specific feature learning in previous research. However, it is difficult for existing methods to address the tradeoff between both issues; e.g., the modality-common feature is learned from the latent commonalities of audio-visual data or the correlated features as aligned projections, in which the modality-specific feature can be lost. To solve the tradeoff, we propose a novel end-to-end architecture, which synchronously projects audio-visual data into the explicit and the implicit dual common subspaces. The explicit subspace is used to learn modality-common features and reduce the modality gap of explicitly paired audio-visual data, where the representation-specific details are abandoned to retain the common underlying structure of audio-visual data. The implicit subspace is used to learn modality-specific features, where each modality privately pulls apart the feature distances between different categories to maintain the category-based distinctions, by minimizing the distance between audio-visual features and corresponding labels. The comprehensive experimental results on two audio-visual datasets, VEGAS and AVE, demonstrate that our proposed model for using two different common subspaces for audio-visual cross-modal learning is effective and significantly outperforms the state-of-the-art cross-modal models that learn features from a single common subspace by 4.30% and 2.30% in terms of average MAP on the VEGAS and AVE datasets, respectively.
Donghuo Zeng, Gen Hattori, Yi Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Complete Cross-triplet Loss in Label Space for Audio-visual Cross-modal Retrieval
abstract
The heterogeneity gap problem is the main challenge in cross-modal retrieval. Because cross-modal data (e.g. audiovisual) have different distributions and representations that cannot be directly compared. To bridge the gap between audiovisual modalities, we learn a common subspace for them by utilizing the intrinsic correlation in the natural synchronization of audio-visual data with the aid of annotated labels. TNN-C-CCA is the best audio-visual cross-modal retrieval (AV-CMR) model so far, but the model training is sensitive to hard negative samples when learning common subspace by applying triplet loss to predict the relative distance between inputs. In this paper, to reduce the interference of hard negative samples in representation learning, we propose a new AV-CMR model to optimize semantic features by directly predicting labels and then measuring the intrinsic correlation between audio-visual data using complete cross-triple loss. In particular, our model projects audio-visual features into label space by minimizing the distance between predicted label features after feature projection and ground label representations. Moreover, we adopt complete cross-triplet loss to optimize the predicted label features by leveraging the relationship between all possible similarity and dissimilarity semantic information across modalities. The extensive experimental results on two audio-visual double-checked datasets have shown an improvement of approximately 2.1% in terms of average MAP over the current state-of-the-art method TNN-C-CCA for the AV-CMR task, which indicates the effectiveness of our proposed model.
Donghuo Zeng, Yanan Wang 0002, Kazushi Ikeda
ISM1
2021 SHECS: A Local Smart Hands-free Elderly Care Support System on Smart AR Glasses with AI Technology
abstract
Some elderly care homes attempt to remedy the shortage of skilled caregivers and provide long-term care for the elderly residents, by enhancing the management of the care support system with the aid of smart devices such as mobile phones and tablets. Since mobile phones and tablets lack the flexibility required for laborious elderly care work, smart AR glasses have already been considered. Although lightweight smart AR devices with a transparent display are more convenient and responsive in an elderly care workplace, fetching data from the server through the Internet results in network congestion not to mention the limited display area. To devise portable smart AR devices that operate smoothly, we first present a no-keepalive-Internet and privacy-compliant required smart hands-free elderly care support system that employs smart glasses with facial recognition and text-to-speech synthesis technologies. Our support system utilizes automatic lightweight facial recognition to identify residents, and information about each resident in question can be obtained hands-free link with a local database. Moreover, a resident information can be displayed immediately on just a portion of the AR smart glasses on the spot. Due to the limited size of the display area, it cannot show all the necessary information. We therefore exploit synthesized voice in the system to read out the elderly care-related information in its entirety. By using the support system, caregivers can gain an understanding of each resident condition immediately, instead of having to devote considerable time in advance in obtaining the complete information of all elderly residents, especially of new residents. Our experiments on this support system were conducted at an elderly care home in Tokyo. Our lightweight facial recognition model achieved high accuracy with fewer model parameters than current state-of-the-art methods. The validation rate of our facial recognition system in the practical use scenario was 99.3% or higher with the false accept rate of 0.001, and caregivers rated the acceptability at 3.6 (5 levels) or higher.
Donghuo Zeng, Tomohiro Obara, Akeri Okawa, Nobuko Iino, Gen Hattori, Ryoichi Kawada, Yasuhiro Takishima
ISM1
2021 TV-watching Companion Robot Supported by Open-domain Chatbot "KACTUS"
abstract
Watching TV once encouraged generations of families and friends [11] to communicate and share empathy. However, the Internet is changing how we watch TV and reducing interaction, leading to problems such as lack of self-control and inadequate communication skills [17]. To understand the conversations while watching TV, we design a scheme based on human conversational behavior [2], and then develop a prototype of TV-watching companion robot supported by the chatbot “KACTUS” [20]. The robot generates a disclosure utterance (e.g., ”I like elephants”) with extracted keywords from the TV program in “TV-watching mode” and uses a cross-topic dialogue management method from “KACTUS” with question utterance to respond with rich conversations in ”Conversation mode”. The robot switches between these two modes at a preset ratio (TV-watching:3, Conversation:1) and behaves like a human enjoying TV-watching. The result of initial experiment shows that three groups of participants enjoyed talking with the robot and the question about their interests in the robot were rated 6.5 (7-levels: ascending from ”extremely disagree” to ”extremely agree”).
Donghuo Zeng, Gen Hattori, Yasuhiro Takishima, Yuta Hagio, Marina Kamimura, Yuta Hoshi, Yutaka Kaneko, Yusei Nishimoto
MUM2
2020 Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to retrieve data in one modality by a query in another modality, which has been a very interesting research issue in the field of multimedia, information retrieval, and computer vision, and database. Most existing works focus on cross-modal retrieval between text-image, text-video, and lyrics-audio. Little research addresses cross-modal retrieval between audio and video due to limited audio-video paired datasets and semantic information. The main challenge of the audio-visual cross-modal retrieval task focuses on learning joint embeddings from a shared subspace for computing the similarity across different modalities, where generating new representations is to maximize the correlation between audio and visual modalities space. In this work, we propose TNN-C-CCA, a novel deep triplet neural network with cluster canonical correlation analysis, which is an end-to-end supervised learning architecture with an audio branch and a video branch. We not only consider the matching pairs in the common space but also compute the mismatching pairs when maximizing the correlation. In particular, two significant contributions are made. First, a better representation by constructing a deep triplet neural network with triplet loss for optimal projections can be generated to maximize correlation in the shared subspace. Second, positive examples and negative examples are used in the learning stage to improve the capability of embedding learning between audio and video. Our experiment is run over fivefold cross validation, where average performance is applied to demonstrate the performance of audio-video cross-modal retrieval. The experimental results achieved on two different audio-visual datasets show that the proposed learning architecture with two branches outperforms existing six canonical correlation analysis–based methods and four state-of-the-art-based cross-modal retrieval methods.
Donghuo Zeng, Yi Yu 0001, Keizo Oyama
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Deep Learning of Human Perception in Audio Event Classification
abstract
In this paper, we introduce our recent studies on human perception in audio event classification. In particular, the pre-trained model VGGish is used as feature extractor to process audio data, and DenseNet is trained by and used as feature extractor for our electroencephalography (EEG) data. The correlation between audio stimuli and EEG is learned in a shared space. In the experiments, we record brain activities (EEG signals) of several subjects while they are listening to music events of 8 audio categories selected from Google AudioSet. Our experimental results demonstrate that i) audio event classification can be improved by exploiting the power of human perception, and ii) the correlation between audio stimuli and EEG can be learned to complement audio event understanding.
Yi Yu 0001, Samuel Beuret, Donghuo Zeng, Keizo Oyama
ISM3
2018 Audio-Visual Embedding for Cross-Modal Music Video Retrieval through Supervised Deep CCA
abstract
Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of different data modalities, such as audio and video, should be taken into account. Music video retrieval by a given musical audio is a natural way to search and interact with music contents. In this work, we study cross-modal music video retrieval in terms of emotion similarity. Particularly, an audio of an arbitrary length is used to retrieve a longer or full-length music video. To this end, we propose a novel audio-visual embedding algorithm by Supervised Deep Canonical Correlation Analysis (S-DCCA) that projects audio and video into a shared space to bridge the semantic gap between audio and video. This also preserves the similarity among audio and visual contents from different videos with the same class label and the temporal structure. The contribution of our approach is mainly manifested in the two aspects: i) We propose to select top k audio chunks by attention-based Long Short-Term Memory (LSTM) model, which can represent good audio summarization with local properties. ii) We propose an end-to-end deep model for crossmodal audio-visual learning where S-DCCA is trained to learn the semantic correlation between audio and visual modalities. Due to the lack of music video dataset, we construct 10K music video dataset from YouTube 8M dataset. Some promising results such as MAP and precision-recall show that our proposed model can be applied to music video retrieval.
Donghuo Zeng, Yi Yu 0001, Keizo Oyama
ISM1
2016 Enlarging drug dictionary with semi-supervised learning for Drug Entity Recognition
abstract
Drug Entity Recognition (DER) is a crucial task for information extraction in biomedical text. Much of previous work for DER using known drugs to build features, however, the known drug resources are limited. In this paper, we proposed a semi-supervised learning to extend an existing drug dictionary. With the extended dictionary, the features for DER can be enriched. Using Conditional Random Fields (CRF) model with the enriched features, an F-measure of 89.26% is achieved on DDIExtraction2013 challenge data set, which outperforms the best system of the DDIExtraction 2013 challenge.
Donghuo Zeng, Chengjie Sun, Lei Lin 0001, Bingquan Liu
BIBM1