Qiuchi Li

dblp:166/3079 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-8219-0869ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Quantum-inspired fuzzy matching network with morphology-enhanced word embeddings
Qiuchi Li, Dawei Song 0001, Prayag Tiwari
Theor. Comput. Sci.2
2026 Adaptive Boosting LLMs for Text Classification
abstract
With large-scale language models demonstrating superior capabilities in a wide range of downstream natural language processing tasks, the future trajectory of research in the field of text categorization faces increasing uncertainty. In this evolving paradigm of open-ended language modeling, where task delimitations are increasingly blurred, a pressing question arises: to what extent has text classification advanced under the full potential of large language model (LLM)? To address this pivotal inquiry, we introduce recurrent generative pre-trained transformer (RGPT), an adaptive boosting framework meticulously designed to craft a dedicated LLM for text classification. RGPT constructs a sequence of base learners by dynamically modulating the training data distribution and iteratively fine-tuning LLMs. These base learners are then progressively integrated, leveraging historical prediction trajectories to form a highly specialized text classification model. Extensive empirical evaluations demonstrate that RGPT surpasses eight state-of-the-art pretrained language models and seven cutting-edge LLMs across four benchmark datasets, achieving an average performance gain of 2.90%.
Yazhou Zhang 0001, Chenyu Ren, Qiuchi Li, Prayag Tiwari, Benyou Wang, Harry Qin
IEEE Trans. Neural Networks Learn. Syst.4
2025 Is Sarcasm Detection a Step-by-Step Reasoning Process in Large Language Models?
abstract
Elaborating a series of intermediate reasoning steps significantly improves the ability of large language models (LLMs) to solve complex problems, as such steps would evoke LLMs to think sequentially. However, human sarcasm understanding is often considered an intuitive and holistic cognitive process, in which various linguistic, contextual, and emotional cues are integrated to form a comprehensive understanding, in a way that does not necessarily follow a step-by-step fashion. To verify the validity of this argument, we introduce a new prompting framework (called SarcasmCue) containing four sub-methods, viz. chain of contradiction (CoC), graph of cues (GoC), bagging of cues (BoC) and tensor of cues (ToC), which elicits LLMs to detect human sarcasm by considering sequential and non-sequential prompting methods. Through a comprehensive empirical comparison on four benchmarks, we highlight three key findings: (1) CoC and GoC show superior performance with more advanced models like GPT-4 and Claude 3.5, with an improvement of 3.5%. (2) ToC significantly outperforms other methods when smaller LLMs are evaluated, boosting the F1 score by 29.7% over the best baseline. (3) Our proposed framework consistently pushes the state-of-the-art (i.e., ToT) by 4.2%, 2.0%, 29.7%, and 58.2% in F1 scores across four datasets. This demonstrates the effectiveness and stability of the proposed framework.
Ben Yao, Yazhou Zhang 0001, Qiuchi Li, Harry Qin
AAAI3
2025 Towards the Law of Capacity Gap in Distilling Language Models
abstract
Language model (LM) distillation aims at distilling the knowledge in a large teacher LM to a small student one. As a critical issue facing LM distillation, a superior student often arises from a teacher of a relatively small scale instead of a larger one, especially in the presence of substantial capacity gap between the teacher and student. This issue, often referred to as the curse of capacity gap, suggests that there is likely an optimal teacher yielding the best-performing student along the scaling course of the teacher. Consequently, distillation trials on teachers of a wide range of scales are called for to determine the optimal teacher, which becomes computationally intensive in the context of large LMs (LLMs). This paper addresses this critical bottleneck by providing the law of capacity gap inducted from a preliminary study on distilling a broad range of small-scale (<3B) LMs, where the optimal teacher consistently scales linearly with the student scale across different model and data scales. By extending the law to LLM distillation on a larger scale (7B), we succeed in obtaining versatile LLMs that outperform a wide array of competitors.
Chen Zhang 0020, Qiuchi Li, Dawei Song 0001, Zheyu Ye, Yan Gao 0017, Yao Hu 0002
ACL (1)2
2025 Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
abstract
Retrieval-augmented generation (RAG) aims to mitigate the hallucination of Large Language Models (LLMs) by retrieving and incorporating relevant external knowledge into the generation process.However, the external knowledge may contain noise and conflict with the parametric knowledge of LLMs, leading to degraded performance.Current LLMs lack inherent mechanisms for resolving such conflicts.To fill this gap, we propose a Dual-Stream Knowledge-Augmented Framework for Shared-Private Semantic Synergy (DSSP-RAG).Central to it is the refinement of the traditional self-attention into a mixed-attention that distinguishes shared and private semantics for a controlled knowledge integration.An unsupervised hallucination detection method that captures the LLMs' intrinsic cognitive uncertainty ensures that external knowledge is introduced only when necessary.To reduce noise in external knowledge, an Energy Quotient (EQ), defined by attention difference matrices between task-aligned and task-misaligned layers, is proposed.Extensive experiments show that DSSP-RAG achieves a superior performance over strong baselines.
Chaozhuo Li, Chen Zhang 0013, Dawei Song 0001, Qiuchi Li
EMNLP5
2025 Learning LLM-as-a-Judge for Preference Alignment
abstract
Learning from preference feedback is a common practice for aligning large language models (LLMs) with human value. Conventionally, preference data is learned and encoded into a scalar reward model that connects a value head with an LLM to produce a scalar score as preference. However, scalar models lack interpretability and are known to be susceptible to biases in datasets. This paper investigates leveraging LLM itself to learn from such preference data and serve as a judge to address both limitations in one shot. Specifically, we prompt the pre-trained LLM to generate initial judgment pairs with contrastive preference in natural language form. The self-generated contrastive judgment pairs are used to train the LLM-as-a-Judge with Direct Preference Optimization (DPO) and incentivize its reasoning capability as a judge. This proposal of learning the LLMas-a-Judge using self-generated Contrastive judgments (Con-J) ensures natural interpretability through the generated rationales supporting the judgments, and demonstrates higher robustness against bias compared to scalar models. Experimental results show that Con-J outperforms the scalar reward model trained on the same collection of preference data, and outperforms a series of open-source and closed-source generative LLMs. We open-source the training process and model weights of Con-J at https://github.com/YeZiyi1998/Con-J.
Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou 0002, Yiqun Liu 0001
ICLR3
2025 Are MLLMs Trapped in the Visual Room?
Yazhou Zhang 0001, Chunwang Zou, Qimeng Liu, Lu Rong, Ben Yao, Zheng Lian 0004, Qiuchi Li, Peng Zhang 0002, Harry Qin
PRCV (7)7
2025 Quantum-inspired semantic matching based on neural networks with the duality of density matrices
abstract
Social media text can be semantically matched in different ways, viz paraphrase identification, answer selection, community question answering, and so on. The performance of the above semantic matching tasks depends largely on the ability of language modeling . Neural network based language models and probabilistic language models are two main streams of language modeling approaches. However, few prior work has managed to unify them in a single framework on the premise of preserving probabilistic features during the neural network learning process. Motivated by recent advances of quantum-inspired neural networks for text representation learning , we fill the gap by resorting to density matrices, a key concept describing a quantum state as well as a quantum probability distribution. The state and probability views of density matrices are mapped respectively to the neural and probabilistic aspects of language models. Concretizing this state-probability duality to the semantic matching task, we build a unified neural-probabilistic language model through a quantum-inspired neural network. Specifically, we take the state view to construct a density matrix representation of sentence, and exploit its probabilistic nature by extracting its main semantics, which form the basis of a legitimate quantum measurement . When matching two sentences, each sentence is measured against the main semantics of the other. Such a process is implemented in a neural structure, facilitating an end-to-end learning of parameters. The learned density matrix representation reflects an authentic probability distribution over the semantic space throughout the training process. Experiments show that our model significantly outperforms a wide range of prominent classical and quantum-inspired baselines.
Qiuchi Li, Dawei Song 0001, Prayag Tiwari
Eng. Appl. Artif. Intell.2
2025 Self-supervised pre-trained neural network for quantum natural language processing
abstract
Quantum computing models have propelled advances in many application domains. However, in the field of natural language processing (NLP), quantum computing models are limited in representation capacity due to the high linearity of the underlying quantum computing architecture. This work attempts to address this limitation by leveraging the concept of self-supervised pre-training, a paradigm that has been propelling the rocketing development of NLP, to increase the power of quantum NLP models on the representation level. Specifically, we present a self-supervised pre-training approach to train quantum encodings of sentences, and fine-tune quantum circuits for downstream tasks on its basis. Experiments show that pre-trained mechanism brings remarkable improvement over end-to-end pure quantum models, yielding meaningful prediction results on a variety of downstream text classification datasets.
Ben Yao, Prayag Tiwari, Qiuchi Li
Neural Networks3
2025 Quantum-inspired neural network with hierarchical entanglement embedding for matching
Zhan Su 0002, Qiuchi Li, Dawei Song 0001, Prayag Tiwari
Neural Networks3
2025 DialogueLLM: Context and emotion knowledge-tuned large language models for emotion recognition in conversations
Yazhou Zhang 0001, Youxi Wu, Prayag Tiwari, Qiuchi Li, Benyou Wang, Harry Qin
Neural Networks5
2024 Task-agnostic Distillation of Encoder-Decoder Language Models
abstract
Finetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from task-specific to task-agnostic distillation of LMs. While task-agnostic, compute-efficient, performance-preserved LMs can be yielded by task-agnostic distillation, previous studies mainly sit in distillation of either encoder-only LMs (e.g., BERT) or decoder-only ones (e.g., GPT) yet largely neglect that distillation of encoder-decoder LMs (e.g., T5) can posit very distinguished behaviors. Frustratingly, we discover that existing task-agnostic distillation methods can fail to handle the distillation of encoder-decoder LMs. To the demand, we explore a few paths and uncover a path named as MiniEnD that successfully tackles the distillation of encoder-decoder LMs in a task-agnostic fashion. We examine MiniEnD on language understanding and abstractive summarization. The results showcase that MiniEnD is generally effective and is competitive compared to other alternatives. We further scale MiniEnD up to distillation of 3B encoder-decoder language models with interpolated distillation. The results imply the opportunities and challenges in distilling large language models (e.g., LLaMA).
Chen Zhang 0020, Yang Yang 0129, Qiuchi Li, Jingang Wang, Dawei Song 0001
LREC/COLING3
2023 Joint Extraction and Classification of Danish Competences for Job Matching
Qiuchi Li, Christina Lioma
ECIR (2)1
2023 CMMA: Benchmarking Multi-Affection Detection in Chinese Multi-Modal Conversations
abstract
Human communication has a multi-modal and multi-affection nature. The inter-relatedness of different emotions and sentiments poses a challenge to jointly detect multiple human affections with multi-modal clues. Recent advances in this field employed multi-task learning paradigms to render the inter-relatedness across tasks, but the scarcity of publicly available resources sets a limit to the potential of works. To fill this gap, we build the first Chinese Multi-modal Multi-Affection conversation (CMMA) dataset, which contains 3,000 multi-party conversations and 21,795 multi-modal utterances collected from various styles of TV-series. CMMA contains a wide variety of affection labels, including sentiment, emotion, sarcasm and humor, as well as the novel inter-correlations values between certain pairs of tasks. Moreover, it provides the topic and speaker information in conversations, which promotes better modeling of conversational context. On the dataset, we empirically analyze the influence of different data modalities and conversational contexts on different affection analysis tasks, and exhibit the practical benefit of inter-task correlations. The full dataset will be publicly available for research\footnote{https://github.com/annoymity2022/Chinese-Dataset}
Yazhou Zhang 0001, Yang Yu 0044, Benyou Wang, Sagar Uprety, Dawei Song 0001, Qiuchi Li, Harry Qin
NeurIPS8
2021 Quantum Cognitively Motivated Decision Fusion for Video Sentiment Analysis
abstract
Video sentiment analysis as a decision-making process is inherently complex, involving the fusion of decisions from multiple modalities and the so-caused cognitive biases. Inspired by recent advances in quantum cognition, we show that the sentiment judgment from one modality could be incompatible with the judgment from another, i.e., the order matters and they cannot be jointly measured to produce a final decision. Thus the cognitive process exhibits ``quantum-like'' biases that cannot be captured by classical probability theories. Accordingly, we propose a fundamentally new, quantum cognitively motivated fusion strategy for predicting sentiment judgments. In particular, we formulate utterances as quantum superposition states of positive and negative sentiment judgments, and uni-modal classifiers as mutually incompatible observables, on a complex-valued Hilbert space with positive-operator valued measures. Experiments on two benchmarking datasets illustrate that our model significantly outperforms various existing decision level and a range of state-of-the-art content-level fusion approaches. The results also show that the concept of incompatibility allows effective handling of all combination patterns, including those extreme cases that are wrongly predicted by all uni-modal classifiers.
Dimitris Gkoumas, Qiuchi Li, Shahram Dehdashti, Massimo Melucci, Yijun Yu 0001, Dawei Song 0001
AAAI2
2021 Quantum-inspired Neural Network for Conversational Emotion Recognition
abstract
We provide a novel perspective on conversational emotion recognition by drawing an analogy between the task and a complete span of quantum measurement. We characterize different steps of quantum measurement in the process of recognizing speakers' emotions in conversation, and stitch them up with a quantum-like neural network. The quantum-like layers are implemented by complex-valued operations to ensure an authentic adoption of quantum concepts, which naturally enables conversational context modeling and multimodal fusion. We borrow an existing algorithm to learn the complex-valued network weights, so that the quantum-like procedure is conducted in a data-driven manner. Our model is comparable to state-of-the-art approaches on two benchmarking datasets, and provide a quantum view to understand conversational emotion recognition.
Qiuchi Li, Dimitris Gkoumas, Alessandro Sordoni, Jian-Yun Nie, Massimo Melucci
AAAI1
2021 An Entanglement-driven Fusion Neural Network for Video Sentiment Analysis
abstract
Video data is multimodal in its nature, where an utterance can involve linguistic, visual and acoustic information. Therefore, a key challenge for video sentiment analysis is how to combine different modalities for sentiment recognition effectively. The latest neural network approaches achieve state-of-the-art performance, but they neglect to a large degree of how humans understand and reason about sentiment states. By contrast, recent advances in quantum probabilistic neural models have achieved comparable performance to the state-of-the-art, yet with better transparency and increased level of interpretability. However, the existing quantum-inspired models treat quantum states as either a classical mixture or as a separable tensor product across modalities, without triggering their interactions in a way that they are correlated or non-separable (i.e., entangled). This means that the current models have not fully exploited the expressive power of quantum probabilities. To fill this gap, we propose a transparent quantum probabilistic neural model. The model induces different modalities to interact in such a way that they may not be separable, encoding crossmodal information in the form of non-classical correlations. Comprehensive evaluation on two benchmarking datasets for video sentiment analysis shows that the model achieves significant performance improvement. We also show that the degree of non-separability between modalities optimizes the post-hoc interpretability.
Dimitris Gkoumas, Qiuchi Li, Yijun Yu 0001, Dawei Song 0001
IJCAI2
2021 CFN: A Complex-Valued Fuzzy Network for Sarcasm Detection in Conversations
abstract
Sarcasm detection in conversation, a theoretically and practically challenging artificial intelligence task, aims to discover elusively ironic, contemptuous, and metaphoric information implied in daily conversations. Most of the recent approaches in sarcasm detection have neglected the intrinsic vagueness and uncertainty of human language in emotional expression and understanding. To address this gap, we propose a complex-valued fuzzy network by leveraging the mathematical formalisms of quantum theory and fuzzy logic. In particular, the target utterance to be recognized is considered as a quantum superposition of a set of separate words. The contextual interaction between adjacent utterances is described as the interaction between a quantum system and its surrounding environment, constructing the quantum composite system, where the weight of interaction is determined by a fuzzy membership function. In order to model both the vagueness and uncertainty, the aforementioned superposition and composite systems are mathematically encapsulated in a density matrix. Finally, a quantum fuzzy measurement is performed on the density matrix of each utterance to yield the probabilistic outcomes of sarcasm recognition. Extensive experiments are conducted on the MUStARD and the 2020 sarcasm detection Reddit track datasets, and the results show that our model outperforms a wide range of strong baselines.
Yazhou Zhang 0001, Yaochen Liu, Qiuchi Li, Prayag Tiwari, Benyou Wang, Hari Mohan Pandey, Peng Zhang 0002, Dawei Song 0001
IEEE Trans. Fuzzy Syst.3
2020 Assessing the Memory Ability of Recurrent Neural Networks
abstract
It is known that Recurrent Neural Networks (RNNs) can remember, in their hidden layers, part of the semantic information expressed by a sequence (e.g., a sentence) that is being processed. Different types of recurrent units have been designed to enable RNNs to remember information over longer time spans. However, the memory abilities of different recurrent units are still theoretically and empirically unclear, thus limiting the development of more effective and explainable RNNs. To tackle the problem, in this paper, we identify and analyze the internal and external factors that affect the memory ability of RNNs, and propose a Semantic Euclidean Space to represent the semantics expressed by a sequence. Based on the Semantic Euclidean Space, a series of evaluation indicators are defined to measure the memory abilities of different recurrent units and analyze their limitations (Code is available at https://github.com/chzhang/Assessing-the-Memory-Ability-of-RNNs). These evaluation indicato...
Cheng Zhang 0019, Qiuchi Li, Lingyu Hua, Dawei Song 0001
ECAI2
2020 Encoding word order in complex embeddings
Benyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li, Peng Zhang 0002, Jakob Grue Simonsen
ICLR4
2019 Aspect-based Sentiment Classification with Aspect-specific Graph Convolutional Networks
abstract
Chen Zhang, Qiuchi Li, Dawei Song. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chen Zhang 0020, Qiuchi Li, Dawei Song 0001
EMNLP/IJCNLP (1)2
2019 Quantum-Inspired Interactive Networks for Conversational Sentiment Analysis
abstract
Conversational sentiment analysis is an emerging, yet challenging Artificial Intelligence (AI) subtask. It aims to discover the affective state of each participant in a conversation. There exists a wealth of interaction information that affects the sentiment of speakers. However, the existing sentiment analysis approaches are insufficient in dealing with this task due to ignoring the interactions and dependency relationships between utterances. In this paper, we aim to address this issue by modeling intrautterance and inter-utterance interaction dynamics. We propose an approach called quantum-inspired interactive networks (QIN), which leverages the mathematical formalism of quantum theory (QT) and the long short term memory (LSTM) network, to learn such interaction dynamics. Specifically, a density matrix based convolutional neural network (DM-CNN) is proposed to capture the interactions within each utterance (i.e., the correlations between words), and a strong-weak influence model inspired by quantum measurement theory is developed to learn the interactions between adjacent utterances (i.e., how one speaker influences another). Extensive experiments are conducted on the MELD and IEMOCAP datasets. The experimental results demonstrate the effectiveness of the QIN model.
Yazhou Zhang 0001, Qiuchi Li, Dawei Song 0001, Peng Zhang 0002
IJCAI2
2019 Multimodal Data Fusion with Quantum Inspiration
abstract
Language understanding is multimodal. During human communication, messages are conveyed not only by words in textual form, but also through speech patterns, gestures or facial emotions of the speakers. Therefore, it is crucial to fuse information from different modalities to achieve a joint comprehension.
Qiuchi Li
SIGIR1
2019 Syntax-Aware Aspect-Level Sentiment Classification with Proximity-Weighted Convolution Network
abstract
It has been widely accepted that Long Short-Term Memory (LSTM) network, coupled with attention mechanism and memory module, is useful for aspect-level sentiment classification. However, existing approaches largely rely on the modelling of semantic relatedness of an aspect with its context words, while to some extent ignore their syntactic dependencies within sentences. Consequently, this may lead to an undesirable result that the aspect attends on contextual words that are descriptive of other aspects. In this paper, we propose a proximity-weighted convolution network to offer an aspect-specific syntax-aware representation of contexts. In particular, two ways of determining proximity weight are explored, namely position proximity and dependency proximity. The representation is primarily abstracted by a bidirectional LSTM architecture and further enhanced by a proximity-weighted convolution. Experiments conducted on the SemEval 2014 benchmark demonstrate the effectiveness of our proposed approach compared with a range of state-of-the-art models is available at https://github.com/GeneZC/PWCN.
Chen Zhang 0020, Qiuchi Li, Dawei Song 0001
SIGIR2
2019 Semantic Hilbert Space for Text Representation Learning
abstract
Capturing the meaning of sentences has long been a challenging task. Current models tend to apply linear combinations of word features to conduct semantic composition for bigger-granularity units e.g. phrases, sentences, and documents. However, the semantic linearity does not always hold in human language. For instance, the meaning of the phrase “ivory tower” cannot be deduced by linearly combining the meanings of “ivory” and “tower”. To address this issue, we propose a new framework that models different levels of semantic units (e.g. sememe, word, sentence, and semantic abstraction) on a single Semantic Hilbert Space, which naturally admits a non-linear semantic composition by means of a complex-valued vector word representation. An end-to-end neural network 1 is proposed to implement the framework in the text classification task, and evaluation results on six benchmarking text classification datasets demonstrate the effectiveness, robustness and self-explanation power of the proposed model. Furthermore, intuitive case studies are conducted to help end users to understand how the framework works.
Benyou Wang, Qiuchi Li, Massimo Melucci, Dawei Song 0001
WWW2
2015 Modeling Multi-query Retrieval Tasks Using Density Matrix Transformation
abstract
The quantum probabilistic framework has recently been applied to Information Retrieval (IR). A representative is the Quantum Language Model (QLM), which is developed for the ad-hoc retrieval with single queries and has achieved significant improvements over traditional language models. In QLM, a density matrix, defined on the quantum probabilistic space, is estimated as a representation of user's search intention with respect to a specific query. However, QLM is unable to capture the dynamics of user's information need in query history. This limitation restricts its further application on the dynamic search tasks, e.g., session search. In this paper, we propose a Session-based Quantum Language Model (SQLM) that deals with multi-query session search task. In SQLM, a transformation model of density matrices is proposed to model the evolution of user's information need in response to the user's interaction with search engine, by incorporating features extracted from both positive feedback (clicked documents) and negative feedback (skipped documents). Extensive experiments conducted on TREC 2013 and 2014 session track data demonstrate the effectiveness of SQLM in comparison with the classic QLM.
Qiuchi Li, Jingfei Li, Peng Zhang 0002, Dawei Song 0001
SIGIR1