EDBT 2026 Demo / reviewers in the wild / expert
Yazhou Zhang 0001
dblp:200/3057-1
· DBLP profile ↗
44ranked-venue papers
20as first author
34since 2021 · last 2026
0000-0002-5699-0176ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Boosting LLMs for Text ClassificationabstractWith large-scale language models demonstrating superior capabilities in a wide range of downstream natural language processing tasks, the future trajectory of research in the field of text categorization faces increasing uncertainty. In this evolving paradigm of open-ended language modeling, where task delimitations are increasingly blurred, a pressing question arises: to what extent has text classification advanced under the full potential of large language model (LLM)? To address this pivotal inquiry, we introduce recurrent generative pre-trained transformer (RGPT), an adaptive boosting framework meticulously designed to craft a dedicated LLM for text classification. RGPT constructs a sequence of base learners by dynamically modulating the training data distribution and iteratively fine-tuning LLMs. These base learners are then progressively integrated, leveraging historical prediction trajectories to form a highly specialized text classification model. Extensive empirical evaluations demonstrate that RGPT surpasses eight state-of-the-art pretrained language models and seven cutting-edge LLMs across four benchmark datasets, achieving an average performance gain of 2.90%. Yazhou Zhang 0001, Chenyu Ren, Qiuchi Li, Prayag Tiwari, Benyou Wang, Harry Qin |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Is Sarcasm Detection a Step-by-Step Reasoning Process in Large Language Models?abstractElaborating a series of intermediate reasoning steps significantly improves the ability of large language models (LLMs) to solve complex problems, as such steps would evoke LLMs to think sequentially. However, human sarcasm understanding is often considered an intuitive and holistic cognitive process, in which various linguistic, contextual, and emotional cues are integrated to form a comprehensive understanding, in a way that does not necessarily follow a step-by-step fashion. To verify the validity of this argument, we introduce a new prompting framework (called SarcasmCue) containing four sub-methods, viz. chain of contradiction (CoC), graph of cues (GoC), bagging of cues (BoC) and tensor of cues (ToC), which elicits LLMs to detect human sarcasm by considering sequential and non-sequential prompting methods. Through a comprehensive empirical comparison on four benchmarks, we highlight three key findings: (1) CoC and GoC show superior performance with more advanced models like GPT-4 and Claude 3.5, with an improvement of 3.5%. (2) ToC significantly outperforms other methods when smaller LLMs are evaluated, boosting the F1 score by 29.7% over the best baseline. (3) Our proposed framework consistently pushes the state-of-the-art (i.e., ToT) by 4.2%, 2.0%, 29.7%, and 58.2% in F1 scores across four datasets. This demonstrates the effectiveness and stability of the proposed framework. Ben Yao, Yazhou Zhang 0001, Qiuchi Li, Harry Qin |
AAAI | 2 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 6 |
| 2025 | Are MLLMs Trapped in the Visual Room?
Yazhou Zhang 0001, Chunwang Zou, Qimeng Liu, Lu Rong, Ben Yao, Zheng Lian 0004, Qiuchi Li, Peng Zhang 0002, Harry Qin |
PRCV (7) | 1 |
| 2025 | BegoniaGPT: Cultivating the large language model to be an exceptional K-12 English teacher
Lu Rong, Yazhou Zhang 0001, Prayag Tiwari |
Neural Networks | 2 |
| 2025 | DialogueLLM: Context and emotion knowledge-tuned large language models for emotion recognition in conversations
Yazhou Zhang 0001, Youxi Wu, Prayag Tiwari, Qiuchi Li, Benyou Wang, Harry Qin |
Neural Networks | 1 |
| 2025 | SarcasmBench: Towards Evaluating Large Language Models on Sarcasm UnderstandingabstractIn the era of large language models (LLMs), tasks associated with “System I” cognition—those that are fast, automatic, and intuitive, such as sentiment analysis and text classification—are often considered effectively solved. However, sarcasm remains a persistent challenge. As a subtle and complex linguistic phenomenon, sarcasm frequently involves rhetorical devices such as hyperbole and figurative language to express implicit sentiments and intentions, demanding a higher level of abstraction and pragmatic reasoning than standard sentiment analysis. This raises concerns about whether current claims of LLM success extend robustly to the domain of sarcasm understanding. To systematically investigate this issue, we introduce a new high-quality multi-modal sarcasm detection dataset, termedAMSD, and construct a comprehensive evaluation benchmark,SarcasmBench. Our benchmark encompasses 16 state-of-the-art (SOTA) LLMs and 8 strong pretrained language models (PLMs), evaluated across six widely-used textual sarcasm datasets and three multi-modal sarcasm benchmarks. We adopt three popular prompting paradigms: zero-shot input/output (IO) prompting, few-shot IO prompting, and chain-of-thought (CoT) prompting. Our extensive experiments yield three key findings: (1) current LLMs underperform supervised PLMs based sarcasm detection baselines. This suggests that significant efforts are still required to improve LLMs' understanding of human sarcasm. (2) GPT-4 and Gemini 2.0 consistently and significantly outperforms other LLMs across various prompting methods. (3) Few-shot IO prompting method outperforms the other two methods: zero-shot IO and few-shot CoT. We hope this benchmark will serve as a valuable resource for the research community and inspire future work toward more robust and human-aligned sarcasm understanding. Yazhou Zhang 0001, Chunwang Zou, Zheng Lian 0004, Prayag Tiwari, Harry Qin |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Chain of Stance: Stance Detection with Large Language Models
Junxia Ma, Changjiang Wang, Hanwen Xing, Yazhou Zhang 0001 |
NLPCC (5) | 5 |
| 2024 | Identification of human microRNA-disease association via low-rank approximation-based link propagation and multiple kernel learning
Yizheng Wang, Xin Zhang 0103, Ying Ju 0002, Quan Zou 0001, Yazhou Zhang 0001, Yijie Ding, Ying Zhang 0060 |
Frontiers Comput. Sci. | 6 |
| 2024 | Learning interactions across sentiment and emotion with graph attention network and position encodings
Ao Jia, Yazhou Zhang 0001, Sagar Uprety, Dawei Song 0001 |
Pattern Recognit. Lett. | 2 |
| 2024 | A Quantum Probability Driven Framework for Joint Multi-Modal Sarcasm, Sentiment and Emotion AnalysisabstractSarcasm, sentiment, and emotion are three typical kinds of spontaneous affective responses of humans to external events and they are tightly intertwined with each other. Such events may be expressed in multiple modalities (e.g., linguistic, visual and acoustic), e.g., multi-modal conversations. Joint analysis of humans’ multi-modal sarcasm, sentiment, and emotion is an important yet challenging topic, as it is a complex cognitive process involving both cross-modality interaction and cross-affection correlation. From the probability theory perspective, cross-affection correlation also means that the judgments on sarcasm, sentiment, and emotion are incompatible. However, this exposed phenomenon cannot be sufficiently modelled by classical probability theory due to its assumption of compatibility. Neither do the existing approaches take it into consideration. In view of the recent success of quantum probability (QP) in modeling human cognition, particularly contextual incompatible decision making, we take the first step towards introducing QP into joint multi-modal sarcasm, sentiment, and emotion analysis. Specifically, we propose aQUantum probabIlity driven multi-modal sarcasm, sEntiment and emoTion analysis framework, termed QUIET. Extensive experiments on two datasets and the results show that the effectiveness and advantages of QUIET in comparison with a wide range of the state-of-the-art baselines. We also show the great potential of QP in multi-affect analysis. Yaochen Liu, Yazhou Zhang 0001, Dawei Song 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | M3GAT: A Multi-modal, Multi-task Interactive Graph Attention Network for Conversational Sentiment Analysis and Emotion RecognitionabstractSentiment and emotion, which correspond to long-term and short-lived human feelings, are closely linked to each other, leading to the fact that sentiment analysis and emotion recognition are also two interdependent tasks in natural language processing (NLP). One task often leverages the shared knowledge from another task and performs better when solved in a joint learning paradigm. Conversational context dependency, multi-modal interaction, and multi-task correlation are three key factors that contribute to this joint paradigm. However, none of the recent approaches have considered them in a unified framework. To fill this gap, we propose a multi-modal, multi-task interactive graph attention network, termed M3GAT, to simultaneously solve the three problems. At the heart of the model is a proposed interactive conversation graph layer containing three core sub-modules, which are: (1) local-global context connection for modeling both local and global conversational context, (2) cross-modal connection for learning multi-modal complementary and (3) cross-task connection for capturing the correlation across two tasks. Comprehensive experiments on three benchmarking datasets, MELD, MEISD, and MSED, show the effectiveness of M3GAT over state-of-the-art baselines with the margin of 1.88%, 5.37%, and 0.19% for sentiment analysis, and 1.99%, 3.65%, and 0.13% for emotion recognition, respectively. In addition, we also show the superiority of multi-task learning over the single-task framework. Yazhou Zhang 0001, Ao Jia, Bo Wang 0011, Peng Zhang 0002, Yuexian Hou, Xiaojia Jin, Dawei Song 0001, Harry Qin |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Self-Adaptive Representation Learning Model for Multi-Modal Sentiment and Sarcasm Joint AnalysisabstractSentiment and sarcasm are intimate and complex, as sarcasm often deliberately elicits an emotional response in order to achieve its specific purpose. Current challenges in multi-modal sentiment and sarcasm joint detection mainly include multi-modal representation fusion and the modeling of the intrinsic relationship between sentiment and sarcasm. To address these challenges, we propose a single-input stream self-adaptive representation learning model (SRLM) for sentiment and sarcasm joint recognition. Specifically, we divide the image into blocks to learn its serialized features and fuse textual feature as input to the target model. Then, we introduce an adaptive representation learning network using a gated network approach for sarcasm and sentiment classification. In this framework, each task is equipped with its dedicated expert network responsible for learning task-specific information, while the shared expert knowledge is acquired and weighted through the gating network. Finally, comprehensive experiments conducted on two publicly available datasets, namely Memotion and MUStARD, demonstrate the effectiveness of the proposed model when compared to state-of-the-art baselines. The results reveal a notable improvement on the performance of sentiment and sarcasm tasks. Yazhou Zhang 0001, Yang Yu 0044, Min Huang 0001, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | MIP-GAT: A Multi-Task Interactive Graph Attention Network with Position Encodings for Joint Sentiment Classification and Emotion Recognition
Ao Jia, Yazhou Zhang 0001, Sagar Uprety, Dawei Song 0001 |
CogSci | 2 |
| 2023 | CMMA: Benchmarking Multi-Affection Detection in Chinese Multi-Modal ConversationsabstractHuman communication has a multi-modal and multi-affection nature. The inter-relatedness of different emotions and sentiments poses a challenge to jointly detect multiple human affections with multi-modal clues. Recent advances in this field employed multi-task learning paradigms to render the inter-relatedness across tasks, but the scarcity of publicly available resources sets a limit to the potential of works. To fill this gap, we build the first Chinese Multi-modal Multi-Affection conversation (CMMA) dataset, which contains 3,000 multi-party conversations and 21,795 multi-modal utterances collected from various styles of TV-series. CMMA contains a wide variety of affection labels, including sentiment, emotion, sarcasm and humor, as well as the novel inter-correlations values between certain pairs of tasks. Moreover, it provides the topic and speaker information in conversations, which promotes better modeling of conversational context. On the dataset, we empirically analyze the influence of different data modalities and conversational contexts on different affection analysis tasks, and exhibit the practical benefit of inter-task correlations. The full dataset will be publicly available for research\footnote{https://github.com/annoymity2022/Chinese-Dataset} Yazhou Zhang 0001, Yang Yu 0044, Benyou Wang, Sagar Uprety, Dawei Song 0001, Qiuchi Li, Harry Qin |
NeurIPS | 1 |
| 2023 | SentiImgBank: A Large Scale Visual Repository for Image Sentiment Analysis
Yazhou Zhang 0001, Lu Rong |
PRCV (11) | 1 |
| 2023 | Multi-user upper limb rehabilitation training system integrating social interaction
Hui Liang 0004, Shiqing Liu, JunJun Pan, Yazhou Zhang 0001, Xiaohang Dong |
Comput. Graph. | 5 |
| 2023 | A Multimodal Coupled Graph Attention Network for Joint Traffic Event Detection and Sentiment ClassificationabstractTraffic events are one of the main causes of traffic accidents, leading to traffic event detection being a challenging research problem in traffic management and intelligent transportation systems (ITSs). The main gap in this task lies in how to extract and represent the valuable information from various kinds of traffic data. Considering the important role that social networks play in traffic data analysis, we argue that sentiment classification and traffic event detection are two closely related tasks in ITSs, where event and sentiment can reveal both explicit and implicit traffic accidents, respectively. Unfortunately, none of the recent approaches in traffic event detection have taken sentiment knowledge into view. This paper proposes a multimodal coupled graph attention network (MCGAT). It aims to construct a multimodal multitask interactive graphical structure where terms (sucha as words, and pixels) are treated as nodes, and their contextual and cross-modal correlations are formalized as edges. The key components are cross-modal and cross-task graph connection layers. The cross-modal graph connection layer captures the multimodal representation, where each node in one modality connects all nodes in another modality. The cross-task graph connection layer is designed by connecting the multimodal node in one task to two single nodes in another task. Empirical evaluation of two benchmarking datasets, such as MGTES and Twitter, shows the effectiveness of the proposed model over state-of-the-art baselines in terms of F1 and accuracy, with significant improvements of 2.4%, 2.4%, 2.7%, and 2.7%. Yazhou Zhang 0001, Prayag Tiwari, Abdulmotaleb El Saddik, M. Shamim Hossain |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Stance-level Sarcasm Detection with BERT and Stance-centered Graph Attention NetworksabstractComputational Linguistics (CL) associated with the Internet of Multimedia Things (IoMT)-enabled multimedia computing applications brings several research challenges, such as real-time speech understanding, deep fake video detection, emotion recognition, home automation, and so on. Due to the emergence of machine translation, CL solutions have increased tremendously for different natural language processing (NLP) applications. Nowadays, NLP-enabled IoMT is essential for its success. Sarcasm detection, a recently emerging artificial intelligence (AI) and NLP task, aims at discovering sarcastic, ironic, and metaphoric information implied in texts that are generated in the IoMT. It has drawn much attention from the AI and IoMT research community. The advance of sarcasm detection and NLP techniques will provide a cost-effective, intelligent way to work together with machine devices and high-level human-to-device interactions. However, existing sarcasm detection approaches neglect the hidden stance behind texts, thus insufficient to exploit the full potential of the task. Indeed, the stance, i.e., whether the author of a text is in favor of, against, or neutral toward the proposition or target talked in the text, largely determines the text’s actual sarcasm orientation. To fill the gap, in this research, we propose a new task: stance-level sarcasm detection (SLSD), where the goal is to uncover the author’s latent stance and based on it to identify the sarcasm polarity expressed in the text. We then propose an integral framework, which consists of Bidirectional Encoder Representations from Transformers (BERT) and a novel stance-centered graph attention networks (SCGAT). Specifically, BERT is used to capture the sentence representation, and SCGAT is designed to capture the stance information on specific target. Extensive experiments are conducted on a Chinese sarcasm sentiment dataset we created and the SemEval-2018 Task 3 English sarcasm dataset. The experimental results prove the effectiveness of the SCGAT framework over state-of-the-art baselines by a large margin. Yazhou Zhang 0001, Dan Ma 0010, Prayag Tiwari, Chen Zhang 0020, Mehedi Masud, Mohammad Shorfuzzaman, Dawei Song 0001 |
ACM Trans. Internet Techn. | 1 |
| 2023 | Metaverse Virtual Social Center for the Elderly Communication During the Social DistancingabstractThe lack of social activities in the elderly for physical reasons can make them feel lonely and prone to depression. With the spread of COVID-19, it is difficult for the elderly to conduct the few social activities stably, causing the elderly to be more lonely. The metaverse is a virtual world that mirrors reality. It allows the elderly to get rid of the constraints of reality and perform social activities stably and continuously, providing new ideas for alleviating the loneliness of the elderly. Through the analysis of the needs of the elderly, a virtual social center framework for the elderly was proposed in this study. Besides, a prototype system was designed according to the framework. The elderly can socialize in virtual reality with metaverse-related technologies and human-computer interaction tools. Additionally, a test was jointly conducted with the chief physician of the geriatric rehabilitation department of a tertiary hospital. The results demonstrated that the mental state of the elderly who had used the virtual social center was significantly better than that of the elderly who had not used it. Thus, virtual social centers alleviated loneliness and depression in older adults. Virtual social centers can help the elderly relieve loneliness and depression when the global epidemic is normalizing and the population is aging. Hence, they have promotion value Hui Liang 0004, Jiupeng Li, JunJun Pan, Yazhou Zhang 0001, Xiaohang Dong |
Virtual Real. Intell. Hardw. | 5 |
| 2022 | Prompt Learning for Multi-modal COVID-19 DiagnosisabstractThe outbreak of COVID-19 pandemic has spread rapidly and severely affected all aspects of human lives. Recent researches has shown artificial intelligence and deep learning based approaches have achieved successful results in detecting diseases. How to accurately and quickly detect COVID-19 has always been the core topic of research. In this paper, we propose a novel approach based on prompt learning for COVID-19 diagnosis. Different from the traditional “pre-training, fine-tuning” paradigm, we propose the prompt-based method that redefine the COVID-19 diagnosis as a masked predict task. Specifically, we adopt an attention mechanism to learn the multi-modal representation of medical image and text, and manually construct a cloze prompt template and a label word set. Selecting the label word corresponding to the maximum probability by pre-training language model. Finally, mapping the prediction results to the disease categories. Experimental results show that our proposed method obtains obvious improvement of 1.2% in terms of Mi-F1 score compared with the state-of-the-art methods. Yang Yu 0044, Lu Rong, Min Huang 0001, Yazhou Zhang 0001, Yijie Ding |
BIBM | 5 |
| 2022 | A Hybrid Model for Depression Detection With Transformer and Bi-directional Long Short-Term MemoryabstractFailure to diagnose and treat depression in a timely manner causes more than three hundred million people suffering from this mental health disorder worldwide. Depression, a global problem, affects not only people’s emotions, but also their physical and mental states. Early detection of depression is very important for the treatment of patients, so we need to achieve excellent accuracy and practicability of depression detection, among which the most important and challenging problem is to design an effective and robust depression detection model. To solve this problem, we propose a hybrid deep learning model, RoBERTa-BiLSTM, to extract features from depression text sequences. We know that the sequence models require a longer computation time as the processing is done sequentially. However, the Transformer models require less execution time with parallelized processing. This model consolidates the strengths of sequence model and Transformer model while suppressing the limitations of sequence model. Specifically, the model maps the words into a compact meaningful word embedding space through the Robustly optimized BERT approach, and then effectively captures the long-distance contextual semantics using the Bidirectional Long Short-Term Memory model. On the DAIC-WOZ and EATD-Corpus benchmark, our experiments demonstrate that our model outperforms state-of-art methods by a substantial margin. Yazhou Zhang 0001, Lu Rong, Yijie Ding |
BIBM | 1 |
| 2022 | Multi-modal Sentiment and Emotion Joint Analysis with a Deep Attentive Multi-task Learning Model
Yazhou Zhang 0001, Lu Rong, Xiang Li 0064 |
ECIR (1) | 1 |
| 2022 | Beyond Emotion: A Multi-Modal Dataset for Human Desire UnderstandingabstractAo Jia, Yu He, Yazhou Zhang, Sagar Uprety, Dawei Song, Christina Lioma. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ao Jia, Yazhou Zhang 0001, Sagar Uprety, Dawei Song 0001, Christina Lioma |
NAACL-HLT | 3 |
| 2022 | A Multibias-Mitigated and Sentiment Knowledge Enriched Transformer for Debiasing in Multimodal Conversational Emotion Recognition
Jinglin Wang, Fang Ma, Yazhou Zhang 0001, Dawei Song 0001 |
NLPCC (1) | 3 |
| 2022 | A fuzzy semantic representation and reasoning model for multiple associative predicates in knowledge graph
Hui Liang 0004, Suzhi Zhang, Yazhou Zhang 0001, Yong Tang 0001 |
Inf. Sci. | 5 |
| 2022 | Affective Interaction: Attentive Representation Learning for Multi-Modal Sentiment ClassificationabstractThe recent booming of artificial intelligence (AI) applications, e.g., affective robots, human-machine interfaces, autonomous vehicles, and so on, has produced a great number of multi-modal records of human communication. Such data often carry latent subjective users’ attitudes and opinions, which provides a practical and feasible path to realize the connection between human emotion and intelligence services. Sentiment and emotion analysis of multi-modal records is of great value to improve the intelligence level of affective services. However, how to find an optimal manner to learn people’s sentiments and emotional representations has been a difficult problem, since both of them involve subtle mind activity. To solve this problem, a lot of approaches have been published, but most of them are insufficient to mine sentiment and emotion, since they have treated sentiment analysis and emotion recognition as two separate tasks. The interaction between them has been neglected, which limits the efficiency of sentiment and emotion representation learning. In this work, emotion is seen as the external expression of sentiment, while sentiment is the essential nature of emotion. We thus argue that they are strongly related to each other where one’s judgment helps the decision of the other. The key challenges are multi-modal fused representation and the interaction between sentiment and emotion. To solve such issues, we design an external knowledge enhanced multi-task representation learning network, termed KAMT. The major elements contain two attention mechanisms, which are inter-modal and inter-task attentions and an external knowledge augmentation layer. The external knowledge augmentation layer is used to extract the vector of the participant’s gender, age, occupation, and of overall color or shape. The main use of inter-modal attention is to capture effective multi-modal fused features. Inter-task attention is designed to model the correlation between sentiment analysis and emotion classification. We perform experiments on three widely used datasets, and the experimental performance proves the effectiveness of the KAMT model. Yazhou Zhang 0001, Prayag Tiwari, Lu Rong, Nojoom A. Alnajem, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Emotion Recognition from Multi-channel EEG Data through A Dual-pipeline Graph Attention NetworkabstractEEG based emotion recognition technology is currently an important concept in artificial intelligence, and also holds great potential in emotional health care. Nevertheless, one major limitation of the prior approaches is they do not capture the relationships between different time-series and channels explicitly, resulting in inevitable low performance, especially in subject-independent recognition settings. In this paper, we propose a novel graph attention network based model to address this issue. Our framework includes dual-pipeline Graph Attention Network layers in parallel to learn the complex dependencies of multi-channel EEG in both temporal and spatial dimensions. The proposed method outperforms other state-of-the-art models on benchmark SEED dataset. Further analysis shows that our method also has good interpretability. As far as we know, it is the first work that introduce graph attention network into EEG based emotion detection research. Xiang Li 0064, Yazhou Zhang 0001, Prayag Tiwari |
BIBM | 3 |
| 2021 | Supercomputer Supported Online Deep Learning Techniques for High Throughput EEG PredictionabstractElectroencephalogram (EEG) is a precise reflection of the brain activities and has been widely studied in clinical medicine, neuroscience, brain interface, etc. Intelligent prediction of the EEG’s evolution accurately plays important roles in several application areas, such as epilepsy seizure forecasting and neonatal brain monitoring. Nevertheless, when the prediction service is deployed as a business service on the Cloud and open for public usage, there are several problems that need to be resolved: (i) how to design a computation platform to process the high-throughput multi-source EEG data, which arrives sequentially and increases rapidly when the services are rapidly promoted, namely tackling the ‘high-throughput computing’ problem; (ii) how to develop a deep learning model to capture the complex EEG distribution as well as the anomaly patterns that could evolve dynamically, namely tackling the ‘concept drift’ problem for non-stationary EEG signals. To tackle these challenges, we propose an Evolutive Convolutional Neural Network (ECNN) and the corresponding supercomputer supported distributed computation system. The ECNN model can dynamically reweighting the sub-structure of the model from data streams in an online learning fashion, by which the capacity scalability and sustainability are introduced into the model. As far as we know, it is the first work that introduce supercomputer supported online deep learning techniques into EEG prediction research. Xiang Li 0064, Yazhou Zhang 0001 |
BIBM | 2 |
| 2021 | Multi-Task Learning for Jointly Detecting Depression and EmotionabstractDepression is a typical mood disease that makes people a persistent feeling of sadness and loss of interest and pleasure. Emotion thus comes into sight and is tightly entangled with depression in that one helps the understanding of the other. Depression and emotion detection has been a new research task. The central challenges in this task are multi-modal interaction and multi-task correlation. The existing approaches treat them as two separate tasks, and fail to model the relationships between them. In this paper, we propose an attentive multi-modal multitask learning framework, called AMM, to generically address such issues. The core modules are two attention mechanisms, viz. inter-modal $(I_{\mathrm{e}})$ and inter-task $(I_{t})$ attentions. The main motivation of $I_{\mathrm{e}}$ attention is to learn multi-modal fused representation. In contrast, Itattention is proposed to learn the relationship between depression detection and emotion recognition. Extensive experiments are conducted on two large scale datasets, i.e., DAIC and multi-modal Getty Image depression (MGID). The results show the effectiveness of the proposed AMM framework, and also shows that AMM obtains better performance for the main task, i.e., depression detection with the help of the secondary emotion recognition task. Yazhou Zhang 0001, Xiang Li 0064, Lu Rong, Prayag Tiwari |
BIBM | 1 |
| 2021 | MedSeq2Seq: A Medical Knowledge Enriched Sequence to Sequence Learning Model for COVID-19 DiagnosisabstractThe COVID-19 pandemic has had a severe impact on humans’ lives and and healthcare systems worldwide. How to early, fastly and accurately diagnose infected patients via multimodal learning is now a research focus. The central challenges in this task mainly lie on multi-modal data representation and multi-modal feature fusion. To solve such challenges, we propose a medical knowledge enriched multi-modal sequence to sequence learning model, termed MedSeq2Seq. The key components include two attention mechanisms, viz. intra-modal (Ia) and inter-model (Ie) attentions, and a medical knowledge augmentation mechanism. The former two mechanisms are to learn multi-modal refined representation, while the latter aims to incorporate external medical knowledge into the proposed model. The experimental results show the effectiveness of the proposed MedSeq2Seq framework over state-of-the-art baselines with a significant improvement of 1%-2%. Yazhou Zhang 0001, Lu Rong, Xiang Li 0064, Prayag Tiwari, Hui Liang 0004 |
BIBM | 1 |
| 2021 | QIRM: A quantum interactive retrieval model for session search
Yuexian Hou, Zhao Li 0007, Yazhou Zhang 0001 |
Neurocomputing | 4 |
| 2021 | Learning interaction dynamics with an interactive LSTM for conversational sentiment analysis
Yazhou Zhang 0001, Prayag Tiwari, Dawei Song 0001, Xiaoliu Mao, Xiang Li 0064, Hari Mohan Pandey |
Neural Networks | 1 |
| 2021 | CFN: A Complex-Valued Fuzzy Network for Sarcasm Detection in ConversationsabstractSarcasm detection in conversation, a theoretically and practically challenging artificial intelligence task, aims to discover elusively ironic, contemptuous, and metaphoric information implied in daily conversations. Most of the recent approaches in sarcasm detection have neglected the intrinsic vagueness and uncertainty of human language in emotional expression and understanding. To address this gap, we propose a complex-valued fuzzy network by leveraging the mathematical formalisms of quantum theory and fuzzy logic. In particular, the target utterance to be recognized is considered as a quantum superposition of a set of separate words. The contextual interaction between adjacent utterances is described as the interaction between a quantum system and its surrounding environment, constructing the quantum composite system, where the weight of interaction is determined by a fuzzy membership function. In order to model both the vagueness and uncertainty, the aforementioned superposition and composite systems are mathematically encapsulated in a density matrix. Finally, a quantum fuzzy measurement is performed on the density matrix of each utterance to yield the probabilistic outcomes of sarcasm recognition. Extensive experiments are conducted on the MUStARD and the 2020 sarcasm detection Reddit track datasets, and the results show that our model outperforms a wide range of strong baselines. Yazhou Zhang 0001, Yaochen Liu, Qiuchi Li, Prayag Tiwari, Benyou Wang, Hari Mohan Pandey, Peng Zhang 0002, Dawei Song 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2019 | Variational Autoencoder based Latent Factor Decoding of Multichannel EEG for Emotion RecognitionabstractRobust cross-subject emotion recognition based on multichannel EEG has always been a hard work. In this work, we hypothesize there exists default brain variables across subjects in emotional processes. Hence, the states of the latent variables that related to emotional processing must contribute to building robust recognition models. We propose to utilize variational autoencoder (VAE) to determine the latent factors from the multichannel EEG. Through sequence modeling method, we examine the emotion recognition performance based on the learnt latent factors. The performance of the proposed methodology is verified on two public datasets (DEAP and SEED), and compared with traditional matrix factorization based (ICA) and autoencoder based (AE) approaches. Experimental results demonstrate that neural network is suitable for unsupervised EEG modeling and our proposed emotion recognition framework achieves the state-of-the-art performance. As far as we know, it is the first work that introduces VAE into multichannel EEG decoding for emotion recognition. Xiang Li 0064, Dawei Song 0001, Yazhou Zhang 0001, Chunyang Niu, Junwei Zhang 0009, Jidong Huo |
BIBM | 4 |
| 2019 | QPIN: A Quantum-inspired Preference Interactive Network for E-commerce RecommendationabstractRecently, recurrent neural networks (RNNs) based methods have achieved profitable performance on mining temporal characteristics in user behavior. However, user preferences are changing over time and have not been fully exploited in e-commerce scenarios. To fill in the gap, we propose an approach, called quantum inspired preference interactive networks (QPIN), which leverages the mathematical formalism of quantum theory (QT) and the long short term memory (LSTM) network, to interactively learn user preferences. Specifically, the tensor product operation is used to model the interaction among a single user's own preferences, i.e. individual preferences. A quantum many-body wave function (QMWF) is employed to model interaction among all users' preferences, i.e. group preferences. Further, we bridge them by deriving a rigorous projection, and thus take the interplay between them into account. Experiments on an Amazon dataset as well as a real-world e-commerce dataset demonstrate the effectiveness of QPIN, which achieves superior performances compared with the state-of-the-art methods in terms of AUC and F1-score. Zhao Li 0007, Yazhou Zhang 0001, Yuexian Hou, Liangzhu Ge |
CIKM | 3 |
| 2019 | Quantum-Inspired DMATT-BiGRU for Conversational Sentiment AnalysisabstractConversational sentiment analysis (CSA) is emergent research field in natural language processing (NLP). This brings a lot of new issues worth studying and directions worth exploring. At the same time, there are many difficulties that need to be overcome. There are lot of challenges, such as lacking of effective deep learning models and being short of appropriate datasets. Inspired by the concept of density matrix in quantum mechanics, we propose a novel attention mechanism called DMATT and apply it to conversational sentiment analysis tasks. In the experiment, we find that deep learning model combined with DMATT has a great improvement in test results compared to the model with traditional attention mechanism. Recurrent neural networks (RNN) and their variants LSTM and GRU are very effective choices in solving time series problem such as conversational sentiment analysis tasks. In this paper, we propose a new model combining GRU and DMATT called DMATT-BiGRU. We experiment in multiple datasets, one of which is called ScenarioSA collected by ourselves. Junwei Zhang 0009, Yuexian Hou, Xiujun Gong, Yazhou Zhang 0001 |
ICTAI | 6 |
| 2019 | Quantum-Inspired Interactive Networks for Conversational Sentiment AnalysisabstractConversational sentiment analysis is an emerging, yet challenging Artificial Intelligence (AI) subtask. It aims to discover the affective state of each participant in a conversation. There exists a wealth of interaction information that affects the sentiment of speakers. However, the existing sentiment analysis approaches are insufficient in dealing with this task due to ignoring the interactions and dependency relationships between utterances. In this paper, we aim to address this issue by modeling intrautterance and inter-utterance interaction dynamics. We propose an approach called quantum-inspired interactive networks (QIN), which leverages the mathematical formalism of quantum theory (QT) and the long short term memory (LSTM) network, to learn such interaction dynamics. Specifically, a density matrix based convolutional neural network (DM-CNN) is proposed to capture the interactions within each utterance (i.e., the correlations between words), and a strong-weak influence model inspired by quantum measurement theory is developed to learn the interactions between adjacent utterances (i.e., how one speaker influences another). Extensive experiments are conducted on the MELD and IEMOCAP datasets. The experimental results demonstrate the effectiveness of the QIN model. Yazhou Zhang 0001, Qiuchi Li, Dawei Song 0001, Peng Zhang 0002 |
IJCAI | 1 |
| 2019 | A quantum-inspired sentiment representation model for twitter sentiment analysis
Yazhou Zhang 0001, Dawei Song 0001, Peng Zhang 0002, Xiang Li 0064 |
Appl. Intell. | 1 |
| 2018 | Unsupervised Sentiment Analysis of Twitter Posts Using Density Matrix Representation
Yazhou Zhang 0001, Dawei Song 0001, Xiang Li 0064, Peng Zhang 0002 |
ECIR | 1 |
| 2018 | Regularizing Deep Neural Networks with an Ensemble-based Decorrelation MethodabstractAlthough Deep Neural Networks (DNNs) have achieved excellent performance in many tasks, improving the generalization capacity of DNNs still remains a challenge. In this work, we propose a novel regularizer named Ensemble-based Decorrelation Method (EDM), which is motivated by the idea of the ensemble learning to improve generalization capacity of DNNs. EDM can be applied to hidden layers in fully connected neural networks or convolutional neural networks. We treat each hidden layer as an ensemble of several base learners through dividing all the hidden units into several non-overlap groups, and each group will be viewed as a base learner. EDM encourages DNNs to learn more diverse representations by minimizing the covariance between all base learners during the training step. Experimental results on MNIST and CIFAR datasets demonstrate that EDM can effectively reduce the overfitting and improve the generalization capacity of DNNs Shuqin Gu, Yuexian Hou, Yazhou Zhang 0001 |
IJCAI | 4 |
| 2018 | Investigating the Dynamic Decision Mechanisms of Users' Relevance Judgment for Information Retrieval via Log Analysis
Jingfei Li, Dawei Song 0001, Pengqing Zhang, Yazhou Zhang 0001 |
PRICAI (1) | 5 |
| 2018 | A quantum-inspired multimodal sentiment analysis frameworkabstractMultimodal sentiment analysis aims to capture diversified sentiment information implied in data that are of different modalities (e.g., an image that is associated with a textual description or a set of textual labels). The key challenge is rooted on the “semantic gap” between different low-level content features and high-level semantic information. Existing approaches generally utilize a combination of multimodal features in a somehow heuristic way. However, how to employ and combine multiple information from different sources effectively is still an important yet largely unsolved problem. To address the problem, in this paper, we propose a Quantum-inspired Multimodal Sentiment Analysis (QMSA) framework. The framework consists of a Quantum-inspired Multimodal Representation (QMR) model (which aims to fill the “semantic gap” and model the correlations between different modalities via density matrix), and a Multimodal decision Fusion strategy inspired by Quantum Interference (QIMF) in the double-slit experiment (in which the sentiment label is analogous to a photon, and the data modalities are analogous to slits). Extensive experiments are conducted on two large scale datasets, which are collected from the Getty Images and Flickr photo sharing platform. The experimental results show that our approach significantly outperforms a wide range of baselines and state-of-the-art methods. Yazhou Zhang 0001, Dawei Song 0001, Peng Zhang 0002, Jingfei Li, Xiang Li 0064, Benyou Wang |
Theor. Comput. Sci. | 1 |
| 2017 | Does tang poetry affect human emotional state? A pilot study by EEGabstractTang poetry, as one of the most typical ways for ancient Chinese to express their emotions, has been continuously inherited for millenniums in China. Nowadays, though our way of life has changed dramatically, the traditional culture about Tang poetry is still affecting us deeply. However, the psychological effect of Tang poetry remains unclear currently. Motivated by this, we aim to investigate the impact of Tang poetry on human emotional state through designing effective experimental paradigm. We gathered 16 Tang poetry video clips with different genres, and played these to 18 subjects, meanwhile recorded their brain neural responses via Electroencephalogram (EEG). A questionnaire was designed to record subjects' subjective ratings of the emotional experience when watching the Tang poetry videos. Through analyzing the questionnaires, we found that Tang poetry can indeed induce the subjects' specific emotions. Finally, we performed a pilot analysis of the recorded EEG signals and the subjective ratings, for exploring the correlation of the brain activity with the emotional states. The results indicate that the Tang poetry can be used as a kind of emotive stimuli in affective computing research, and the EEG induced by Tang poetry can be utilized to probe the human's internal emotions. Yazhou Zhang 0001, Xiang Li 0064, Yuexian Hou, Dawei Song 0001 |
BIBM | 2 |