Jinming Zhao

dblp:121/8902 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Boundary-aware shape recognition using dynamic graph convolutional networks
Jinming Zhao, Junyu Dong, Huiyu Zhou 0001, Xinghui Dong
Pattern Recognit.1
2026 CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations
Zisen Zhou, Chang Wen, Xuefei Liu, Jianhua Tao 0001, Zhengqi Wen, Zheng Lian 0004, Jinming Zhao, Bingsen Xiong, Shaozheng Qin
Pattern Recognit.9
2025 Scaling A Simple Approach to Zero-Shot Speech Recognition
abstract
Despite rapid progress in increasing the language coverage of automatic speech recognition, the field is still far from covering all languages with a known writing script. Recent work showed promising results with a zero-shot approach requiring only a small amount of text data, however, accuracy heavily depends on the quality of the used phonemizer which is often weak for unseen languages. In this paper, we present MMS Zero-shot, a conceptually simpler approach based on romanization and an acoustic model trained on data in 1,078 different languages or three orders of magnitude more than prior art. MMS Zero-shot reduces the average character error rate by a relative 46% over 100 unseen languages compared to the best previous work. Moreover, the error rate of our approach is only 2.5x higher than in-domain supervised baselines, while MMS Zero-shot uses no labeled data for the evaluation languages at all.
Jinming Zhao, Vineel Pratap, Michael Auli
ICASSP1
2025 Continual Speech Learning with Fused Speech Features
Guitao Wang, Jinming Zhao, Guilin Qi, Tongtong Wu, Gholamreza Haffari
INTERSPEECH2
2025 Exploring Interpretability in Deep Learning for Affective Computing: A Comprehensive Review
abstract
Deep learning has shown impressive performance in affective computing, but its black-box characteristic limits the model’s interpretability, posing a challenge to further development and application. Compared with objective recognition tasks such as image recognition, emotion perception as a high-level cognition is more subjective, making it particularly important to enhance the interpretability of deep learning in affective computing. In recent years, some interpretability-related works have emerged, but there are few reviews on this topic yet. This article summarizes the explainable deep learning methods in affective computing from two aspects: first, the application of general explainable deep learning methods in affective computing from the perspectives of model-agnostic and model-specific is introduced; second, emotion-specific interpretability research that combines emotional psychology theories, physiological studies, and human cognition, covering task design, model design, and result analysis methods, is systematically reviewed. There are new explainable deep learning methods for multimodal and large language models in the context of emotion. Finally, we discuss five specific challenges and propose corresponding future directions to provide insights and references for subsequent research on affective computing interpretability.
Tenggan Zhang, Jinming Zhao, Qin Jin
ACM Trans. Multim. Comput. Commun. Appl.4
2024 ESCoT: Towards Interpretable Emotional Support Dialogue Systems
abstract
Understanding the reason for emotional support response is crucial for establishing connections between users and emotional support dialogue systems.Previous works mostly focus on generating better responses but ignore interpretability, which is extremely important for constructing reliable dialogue systems.To empower the system with better interpretability, we propose an emotional support response generation scheme, named Emotion-Focused and Strategy-Driven Chain-of-Thought (ESCoT), mimicking the process of identifying, understanding, and regulating emotions.Specially, we construct a new dataset with ESCoT in two steps: (1) Dialogue Generation where we first generate diverse conversation situations, then enhance dialogue generation using richer emotional support strategies based on these situations; (2) Chain Supplement where we focus on supplementing selected dialogues with elements such as emotion, stimulus, appraisal, and strategy reason, forming the manually verified chains.Additionally, we further develop a model to generate dialogue responses with better interpretability.We also conduct extensive experiments and human evaluations to validate the effectiveness of the proposed ESCoT and generated dialogue responses.Our data and code are available at https://github.com/TeigenZhang/ESCoT.
Tenggan Zhang, Jinming Zhao, Qin Jin
ACL (1)3
2024 NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation Corpus
abstract
It remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligned parallel English-Japanese SI dataset. Starting with a non-aligned corpus NAIST-SIC, we propose a two-stage alignment approach to make the corpus parallel and thus suitable for model training. The first stage is coarse alignment where we perform a many-to-many mapping between source and target sentences, and the second stage is fine-grained alignment where we perform intra- and inter-sentence filtering to improve the quality of aligned pairs. To ensure the quality of the corpus, each step has been validated either quantitatively or qualitatively. This is the first open-sourced large-scale parallel SI dataset in the literature. We also manually curated a small test set for evaluation purposes. Our results show that models trained with SI data lead to significant improvement in translation quality and latency over baselines. We hope our work advances research on SI corpora construction and SiMT. Our data will be released upon the paper’s acceptance.
Jinming Zhao, Katsuhito Sudoh, Satoshi Nakamura 0001, Yuka Ko, Kosuke Doi, Ryo Fukuda
LREC/COLING1
2024 Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues
abstract
Personality recognition aims to identify the personality traits implied in user data such as dialogues and social media posts.Current research predominantly treats personality recognition as a classification task, failing to reveal the supporting evidence for the recognized personality.In this paper, we propose a novel task named Explainable Personality Recognition, aiming to reveal the reasoning process as supporting evidence of the personality trait.Inspired by personality theories, personality traits are made up of stable patterns of personality state, where the states are short-term characteristic patterns of thoughts, feelings, and behaviors in a concrete situation at a specific moment in time.We propose an explainable personality recognition framework called Chainof-Personality-Evidence (CoPE), which involves a reasoning process from specific contexts to short-term personality states to longterm personality traits.Furthermore, based on the CoPE framework, we construct an explainable personality recognition dataset from dialogues, PersonalityEvd.We introduce two explainable personality state recognition and explainable personality trait recognition tasks, which require models to recognize the personality state and trait labels and their corresponding support evidence.Our extensive experiments based on Large Language Models on the two tasks show that revealing personality traits is very challenging and we present some insights for future research.
Jinming Zhao, Qin Jin
EMNLP2
2024 ECR-Chain: Advancing Generative Language Models to Better Emotion-Cause Reasoners through Reasoning Chains
Zhaopei Huang, Jinming Zhao, Qin Jin
IJCAI2
2023 Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing Modalities
abstract
Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing data across modalities, the inherent difference between heterogeneous modalities, namely the modality gap, presents a challenge. To address this, we propose to use invariant features for a missing modality imagination network (IF-MMIN) which includes two novel mechanisms: 1) an invariant feature learning strategy that is based on the central moment discrepancy (CMD) distance under the full-modality scenario; 2) an invariant feature based imagination module (IF-IM) to alleviate the modality gap during the missing modalities prediction, thus improving the robustness of multimodal joint representation. Comprehensive experiments on the benchmark dataset IEMOCAP demonstrate that the proposed model outperforms all baselines and invariantly improves the overall emotion recognition performance under uncertain missing-modality conditions. We release the code at: https://github.com/ZhuoYulang/IF-MMIN.
Haolin Zuo, Rui Liu 0008, Jinming Zhao, Guanglai Gao, Haizhou Li 0001
ICASSP3
2023 Investigating Pre-trained Audio Encoders in the Low-Resource Condition
abstract
Pre-trained speech encoders have been central to pushing state-of-the-art results across various speech understanding and generation tasks. Nonetheless, the capabilities of these encoders in low-resource settings are yet to be thoroughly explored. To address this, we conduct a comprehensive set of experiments using a representative set of 3 state-of-the-art encoders (Wav2vec2, WavLM, Whisper) in the low-resource setting across 7 speech understanding and generation tasks. We provide various quantitative and qualitative analyses on task performance, convergence speed, and representational properties of the encoders. We observe a connection between the pre-training protocols of these encoders and the way in which they capture information in their internal layers. In particular, we observe the Whisper encoder exhibits the greatest low-resource capabilities on content-driven tasks in terms of performance and convergence speed.
Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi
INTERSPEECH2
2023 MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
abstract
The first Multimodal Emotion Recognition Challenge (MER 2023)1 was successfully held at ACM Multimedia. The challenge focuses on system robustness and consists of three distinct tracks: (1) MER-MULTI, where participants are required to recognize both discrete and dimensional emotions; (2) MER-NOISE, in which noise is added to test videos for modality robustness evaluation; (3) MER-SEMI, which provides a large amount of unlabeled samples for semi-supervised learning. In this paper, we introduce the motivation behind this challenge, describe the benchmark dataset, and provide some statistics about participants. To continue using this dataset after MER 2023, please sign a new End User License Agreement2 and send it to our official email address3. We believe this high-quality dataset can become a new benchmark in multimodal emotion recognition, especially for the Chinese research community.
Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Jinming Zhao, Ye Liu 0010, Bin Liu 0041, Jiangyan Yi, Meng Wang 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia10
2023 Two-Stage Adaptation for Cross-Corpus Multimodal Emotion Recognition
Zhaopei Huang, Jinming Zhao, Qin Jin
NLPCC (2)2
2022 M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database
abstract
Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, Haizhou Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Jinming Zhao, Tenggan Zhang, Jingwen Hu 0003, Yuchen Liu 0003, Qin Jin, Xinchao Wang, Haizhou Li 0001
ACL (1)1
2022 DialogueEIN: Emotion Interaction Network for Dialogue Affective Analysis
abstract
Emotion Recognition in Conversation (ERC) has attracted increasing attention in the affective computing research field. Previous works have mainly focused on modeling the semantic interactions in the dialogue and implicitly inferring the evolution of the speakers’ emotional states. Few works have considered the emotional interactions, which directly reflect the emotional evolution of speakers in the dialogue. According to psychological and behavioral studies, the emotional inertia and emotional stimulus are important factors that affect the speaker’s emotional state in conversations. In this work, we propose a novel Dialogue Emotion Interaction Network, DialogueEIN, to explicitly model the intra-speaker, inter-speaker, global and local emotional interactions to respectively simulate the emotional inertia, emotional stimulus, global and local emotional evolution in dialogues. Extensive experiments on four ERC benchmark datasets, IEMOCAP, MELD, EmoryNLP and DailyDialog, show that our proposed DialogueEIN considering emotional interaction factors can achieve superior or competitive performance compared to state-of-the-art methods. Our codes and models are released.
Yuchen Liu 0003, Jinming Zhao, Jingwen Hu 0003, Qin Jin
COLING2
2022 Towards relation extraction from speech
abstract
Relation extraction has focused on extracting semantic relationships between entities from the unstructured written textual data.However, with the vast and rapidly increasing amounts of spoken data, relation extraction from speech is an important but under-explored problem. In this paper, we propose a new information extraction task, speech relation extraction (SpeechRE).To facilitate further research, we construct the first synthetic training datasets, as well as the first human-spoken test set with native English speakers.We establish strong baseline performance for SpeechRE via two approaches.The pipeline approach connects a pretrained ASR module with a text-based relation extraction module.The end-to-end approach employs a cross-modal encoder-decoder architecture.Our comprehensive experiments reveal the relative strengths and weaknesses of these approaches, and shed light on important future directions in SpeechRE research.We share the source code and datasets on https://github.com/ wutong8023/SpeechRE.* denotes the equal contribution.
Tongtong Wu, Guitao Wang, Jinming Zhao, Zhaoran Liu, Guilin Qi, Yuan-Fang Li, Gholamreza Haffari
EMNLP3
2022 Memobert: Pre-Training Model with Prompt-Based Learning for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity. In this paper, we propose a multimodal pre-training model MEmoBERT for multimodal emotion recognition, which learns multimodal joint representations through self-supervised learning from a self-collected large-scale unlabeled video data that come in sheer volume. Furthermore, unlike the conventional "pre-train, finetune" paradigm, we propose a prompt-based method that reformulates the downstream emotion classification task as a masked text prediction one, bringing the downstream task closer to the pre-training. Extensive experiments on two benchmark datasets, IEMOCAP and MSP-IMPROV, show that our proposed MEmoBERT significantly enhances emotion recognition performance.
Jinming Zhao, Qin Jin, Xinchao Wang, Haizhou Li 0001
ICASSP1
2022 M-Adapter: Modality Adaptation for End-to-End Speech-to-Text Translation
abstract
End-to-end speech-to-text translation models are often initialized with pre-trained speech encoder and pre-trained text decoder. This leads to a significant training gap between pretraining and fine-tuning, largely due to the modality differences between speech outputs from the encoder and text inputs to the decoder. In this work, we aim to bridge the modality gap between speech and text to improve translation quality. We propose M-Adapter, a novel Transformer-based module, to adapt speech representations to text. While shrinking the speech sequence, M-Adapter produces features desired for speech-to-text translation via modelling global and local dependencies of a speech sequence. Our experimental results show that our model outperforms a strong baseline by up to 1 BLEU score on the Must-C En→DE dataset.
Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi
INTERSPEECH1
2022 Energy Saving in LEO-HTS Constellation Based on Adaptive Power Allocation with Multi-Beam Directivity Control
abstract
High throughput satellite (HTS) has been identified as a key technology for improving communication capacity of satellite system. Recently, low earth orbit high throughput satellite (LEO-HTS) has attracted much attention because of its low latency and construction cost. However, LEO-HTS is power-limited when it passes through the shadow of the earth, which will affect the lifetime and capabilities of the satellite. This paper focuses on the problem of energy saving for multiple LEO-HTSs in the LEO-HTS constellation. A method of adaptive power allocation with multi-beam directivity control according to the traffic demand of user equipments (UEs) is proposed to save energy. In this case, an optimization model is established to minimize the transmission power of multiple LEO-HTSs. Then, an alternating optimization framework is introduced to iteratively solve the transmission power and multi-beam directivity of multiple LEO-HTSs. Finally, numerical results demonstrate that the proposed method can save energy more effectively than the traditional method in the scenario with high traffic demand and concentrated UE distribution.
Jinming Zhao, Yong Li 0001, Zeyu Hu
PIMRC1
2022 Interference Coordination Method for Integrated HAPS-Terrestrial Networks
abstract
Non-terrestrial network (NTN) is an important technique to provide extreme coverage towards 5G advanced and 6G. High-altitude-platform-station (HAPS), as one key factor in NTN, is attracting a lot of attentions. To achieve better resource utilization and performance, a unified design for integrated HAPS-Terrestrial networks is necessary, where the interference among these two systems may be important. Current methods usually consider fixed resource allocation among the two systems without considering the distribution of traffic load and hence have low resource utilization. In practice, the traffic load may change dynamically or semi-statically due to the wide coverage in HAPS, which should be considered in the design of integrated HAPS-terrestrial systems. In this paper, we proposed an interference coordination method based on the distribution of traffic load as well as the deployment of HAPS and terrestrial networks. The evaluation results show that the proposed method achieves higher throughput than existing fixed resource allocation method.
Wenjia Liu, Xiaolin Hou, Lan Chen 0004, Yuki Hokazono, Jinming Zhao
VTC Spring5
2021 MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation
abstract
Jingwen Hu, Yuchen Liu, Jinming Zhao, Qin Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jingwen Hu 0003, Yuchen Liu 0003, Jinming Zhao, Qin Jin
ACL/IJCNLP (1)3
2021 Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities
abstract
Jinming Zhao, Ruichen Li, Qin Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jinming Zhao, Qin Jin
ACL/IJCNLP (1)1
2021 It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation Data
abstract
Most existing simultaneous machine translation (SiMT) systems are trained and evaluated on offline translation corpora.We argue that SiMT systems should be trained and tested on real interpretation data.To illustrate this argument, we propose an interpretation test set and conduct a realistic evaluation of SiMT trained on offline translations.Our results, on our test set along with 3 existing smaller scale language pairs, highlight the difference of up-to 13.83 BLEU score when SiMT models are evaluated on translation vs interpretation data.In the absence of interpretation training data, we propose a translationto-interpretation (T2I) style transfer method which allows converting existing offline translations into interpretation-style data, leading to up-to 2.8 BLEU improvement.However, the evaluation gap remains notable, calling for constructing large-scale interpretation corpora better suited for evaluating and developing SiMT systems. 1
Jinming Zhao, Philip Arthur, Gholamreza Haffari, Trevor Cohn, Ehsan Shareghi
EMNLP (1)1
2021 Speech Emotion Recognition via Multi-Level Cross-Modal Distillation
Jinming Zhao, Qin Jin
Interspeech2
2020 SummPip: Unsupervised Multi-Document Summarization with Sentence Graph Compression
abstract
Obtaining training data for multi-document Summarization (MDS) is time consuming and resource-intensive, so recent neural models can only be trained for limited domains. In this paper, we propose SummPip: an unsupervised method for multi-document summarization, in which we convert the original documents to a sentence graph, taking both linguistic and deep representation into account, then apply spectral clustering to obtain multiple clusters of sentences, and finally compress each cluster to generate the final summary. Experiments on Multi-News and DUC-2004 datasets show that our method is competitive to previous unsupervised methods and is even comparable to the neural supervised approaches. In addition, human evaluation shows our system produces consistent and complete summaries compared to human written ones.
Jinming Zhao, Ming Liu 0028, Longxiang Gao, Lan Du 0002, He Zhao 0001, He Zhang 0034, Gholamreza Haffari
SIGIR1
2019 Cross-culture Multimodal Emotion Recognition with Adversarial Learning
abstract
With the development of globalization, automatic emotion recognition has faced a new challenge in the multi-culture scenario - to generalize across different cultures. Previous works mainly rely on multi-cultural datasets to address the cross-culture discrepancy, which are expensive to collect. In this paper, we propose an adversarial learning framework to alleviate the culture influence on multimodal emotion recognition. We treat the emotion recognition and culture recognition as two adversarial tasks. The emotion feature embedding is trained to improve the emotion recognition but to confuse the culture recognition, so that it is more emotion-salient and culture-invariant for cross-culture emotion recognition. Our approach is applicable to both mono-culture and multi-culture emotion datasets. Extensive experiments demonstrate that the proposed method significantly outperforms previous baselines in both cross-culture and multi-culture evaluations.
Jingjun Liang, Shizhe Chen, Jinming Zhao, Qin Jin
ICASSP3
2019 Speech Emotion Recognition in Dyadic Dialogues with Attentive Interaction Modeling
Jinming Zhao, Shizhe Chen, Jingjun Liang, Qin Jin
INTERSPEECH1
2017 Emotion recognition with multimodal features and temporal models
abstract
This paper presents our methods to the Audio-Video Based Emotion Recognition subtask in the 2017 Emotion Recognition in the Wild (EmotiW) Challenge. The task aims to predict one of the seven basic emotions for short video segments. We extract different features from audio and facial expression modalities. We also explore the temporal LSTM model with the input of frame facial features, which improves the performance of the non-temporal model. The fusion of different modality features and the temporal model lead us to achieve a 58.5% accuracy on the testing set, which shows the effectiveness of our methods.
Wenxuan Wang 0001, Jinming Zhao, Shizhe Chen, Qin Jin, Shilei Zhang, Yong Qin 0001
ICMI3
2017 Deep neural network bottleneck features for bird species verification
abstract
Recently, bottleneck features as effective representations have been successfully used in Speaker Recognition (SR) and Language Recognition (LR), but little work has focused on bottleneck features for Bird Species Verification (BSV). In SR, LR and BSR tasks, using short-time spectra features may be insufficient, so it need some more abstract and discriminative representations as complementation to conventional spectra features. Some SR and LR work shows that bottleneck features can form a low-dimension representation of the original inputs with a powerful descriptive and discriminative capability. Due to the general audio representation principles of speakers, language and birds being similar, we propose a hypothesis: the bottleneck features are also useful for BSV. Therefore, in this paper, we use the bottleneck feature framework based on the standard i-vector framework to deal with crucial problems in conventional methods of BSV, such as the session variability and insufficient features. Moreover, we make no distinction between bird calls and bird songs in the evaluation phase. Experimental results show that the standard i-vector system and the bottleneck feature system gain 3.39% and 0.85% Equal Error Rate (EER) respectively. The bottleneck feature system obtains 75% relative improvement over the standard i-vector system, meaning that the bottleneck features as a complementation to spectra features are significantly useful for BSV. The deep feature system, which is an another state-of-the-art framework based on deep features used in SR, however, only results in 18.64% EER, which is much worse than the other two systems, and a brief explanation is provided in this paper.
Jinming Zhao, Yanyan Xu 0001, Dengfeng Ke, Kaile Su
IJCNN1