VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Quan
dblp:90/5936
· DBLP profile ↗
73ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0002-8385-1083ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 5 first-author · 39 since 2021Databases, data management, data science and information retrieval · 13 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 9 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProFuser: Progressive Fusion of Large Language ModelsabstractWhile fusing the capacities and advantages of various large language models offers a pathway to construct more powerful and versatile models, a fundamental challenge is to properly select advantageous model during training. Existing fusion methods primarily focus on the training mode that uses cross entropy on ground truth in a teacher-forcing setup to measure a model's advantage, which may provide limited insight towards model advantage. In this paper, we introduce a novel approach that enhances the fusion process by incorporating both the training and inference modes. Our method evaluates model advantage not only through cross entropy during training but also by considering inference outputs, providing a more comprehensive assessment. To combine the two modes effectively, we introduce ProFuser to progressively transition from inference mode to training mode. To validate ProFuser's effectiveness, we fused three models, including Vicuna-7B-v1.5, Llama-2-7B-Chat, and MPT-7B-8K-Chat, and demonstrated the improved performance in knowledge, reasoning, and safety compared to baseline methods. Tianyuan Shi, Fanqi Wan, Canbin Huang, Xiaojun Quan, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Minhua Huang 0002, Wu Kai |
AAAI | 4 |
| 2026 | ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue AgentsabstractProactive dialogue has emerged as a critical and challenging research problem in advancing large language models (LLMs).Existing works predominantly focus on domain-specific or task-oriented scenarios, which leads to fragmented evaluations and limits the comprehensive exploration of models' proactive dialogue abilities.In this work, we propose Proac-tiveEval, a unified framework for evaluating proactive dialogue capabilities of LLMs.This framework decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains.Moreover, it also enables the automatic generation of diverse and challenging evaluation data.Based on the proposed framework, we develop 328 evaluation environments spanning 6 distinct domains.Through experiments with 22 different types of LLMs, we show that DeepSeek-R1 and Claude-3.7-Sonnetexhibit exceptional performance on target planning and dialogue guidance tasks, respectively.Finally, we investigate how reasoning capabilities influence proactive behaviors and discuss their implications for future model development.Our code and data are available at the repository. Fanqi Wan, Jiajian Guo, Xiaojun Quan |
ACL (1) | 4 |
| 2025 | Cool-Fusion: Fuse Large Language Models without TrainingabstractWe focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths.One of the challenges of model fusion is high computational load, specifically in fine-tuning or aligning vocabularies.To address this, we propose Cool-Fusion, a simple yet effective approach that fuses the knowledge of source LLMs, which does not require training.Unlike ensemble methods, Cool-Fusion is applicable to any set of source LLMs that have different vocabularies.To overcome the vocabulary discrepancies among LLMs, we ensemble LLMs on text level, allowing them to rerank the generated texts by each other with different granularities.Extensive experiments have been conducted across a variety of benchmark datasets.On GSM8K, Cool-Fusion increases accuracy from three strong source LLMs by a significant margin of 17.4%. Cong Liu 0001, Xiaojun Quan, Yan Pan 0002, Weigang Wu, Xu Chen 0004, Liang Lin 0004 |
ACL (1) | 2 |
| 2025 | Mutual-Taught for Co-adapting Policy and Reward ModelsabstractTianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Ziyi Yang, Weizhou Shen, Xiaojun Quan, Ming Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Weizhou Shen, Xiaojun Quan, Ming Yan 0008 |
ACL (1) | 7 |
| 2025 | Edit-Wise Preference Optimization for Grammatical Error CorrectionabstractWhile large language models (LLMs) have achieved remarkable success in various natural language processing tasks, their strengths have yet to be fully demonstrated in grammatical error correction (GEC). This is partly due to the misalignment between their pre-training objectives and the GEC principle of making minimal edits. In this work, we aim to bridge this gap by introducing a novel method called Edit-wise Preference Optimization (EPO). By distinguishing the importance of different tokens and assigning higher reward weights to edit tokens during preference optimization, our method captures fine-grained distinctions in GEC that traditional preference learning often overlooks. Extensive experiments on both English and Chinese datasets show that our framework consistently outperforms strong baselines, achieving state-of-the-art performance and demonstrating the advantages of LLMs in GEC. Jiehao Liang, Haihui Yang, Shiping Gao, Xiaojun Quan |
COLING | 4 |
| 2025 | FuseChat: Knowledge Fusion of Chat ModelsabstractWhile training large language models (LLMs) from scratch can indeed lead to models with distinct capabilities and strengths, it incurs substantial costs and may lead to redundancy in competencies.Knowledge fusion aims to integrate existing LLMs of diverse architectures and capabilities into a more potent LLM through lightweight continual training, thereby reducing the need for costly LLM development.In this work, we propose a new framework for the knowledge fusion of chat LLMs through two main stages, resulting in FUSECHAT.Firstly, we conduct pairwise knowledge fusion on source chat LLMs of varying structures and scales to create multiple target LLMs with identical structure and size via lightweight fine-tuning.During this process, a statistics-based token alignment approach is introduced as the cornerstone for fusing LLMs with different structures.Secondly, we merge these target LLMs within the parameter space, where we propose a novel method for determining the merging coefficients based on the magnitude of parameter updates before and after fine-tuning.We implement and validate FUSECHAT using six prominent chat LLMs with diverse architectures and scales.Experimental results on two instruction-following benchmarks, AlpacaEval 2.0 and MT-Bench, demonstrate the superiority of FUSECHAT-7B over baselines of various sizes. Fanqi Wan, Longguang Zhong, Ruijun Chen 0001, Xiaojun Quan |
EMNLP | 5 |
| 2025 | Advantage-Guided Distillation for Preference Alignment in Small Language ModelsabstractAlignment techniques enable Large Language Models (LLMs) to generate outputs that align with human preferences and play a crucial role in their effectiveness. However, their impact often diminishes when applied to Small Language Models (SLMs), likely due to the limited capacity of these models. Instead of directly applying existing alignment techniques to SLMs, we propose to utilize a well-aligned teacher LLM to guide the alignment process for these models, thereby facilitating the transfer of the teacher's knowledge of human preferences to the student model. To achieve this, we first explore a straightforward approach, Dual-Constrained Knowledge Distillation (DCKD), that employs knowledge distillation with two KL-divergence constraints from the aligned teacher to the unaligned student. To further enhance the student's ability to distinguish between preferred and dispreferred responses, we then propose Advantage-Guided Distillation for Preference Alignment (ADPA), which leverages an advantage function from the aligned teacher to deliver more nuanced, distribution-level reward signals for the student's alignment. Our experimental results show that these two approaches appreciably improve the alignment of SLMs and narrow the performance gap with larger counterparts. Among them, ADPA demonstrates superior performance and achieves even greater effectiveness when integrated with DCKD. Our code is available at https://github.com/SLIT-AI/ADPA . Shiping Gao, Fanqi Wan, Jiajian Guo, Xiaojun Quan, Qifan Wang 0001 |
ICLR | 4 |
| 2025 | Weighted-Reward Preference Optimization for Implicit Model FusionabstractWhile fusing heterogeneous open-source LLMs with varying architectures and sizes can potentially integrate the strengths of different models, existing fusion methods face significant challenges, such as vocabulary alignment and merging distribution matrices. These procedures are not only complex but also prone to introducing noise and errors. In this paper, we propose an implicit fusion method, Weighted-Reward Preference Optimization (WRPO), which leverages preference optimization between the source LLMs and the target LLM to transfer their capabilities effectively. WRPO eliminates the need for vocabulary alignment and matrix fusion and can be efficiently scaled to accommodate various LLMs. To address distributional deviations between the source and target LLMs, WRPO introduces a progressive adaptation strategy that gradually shifts reliance on preferred examples from the target LLM to the source LLMs. Extensive experiments on the MT-Bench, AlpacaEval-2, and Arena-Hard benchmarks demonstrate that WRPO consistently outperforms existing knowledge fusion methods and various fine-tuning baselines. When applied to LLaMA3-8B-Instruct as the target model, WRPO achieves a length-controlled win rate of 55.9\% against GPT-4-Preview-1106 on AlpacaEval-2 and a win rate of 46.2\% against GPT-4-0314 on Arena-Hard. Our code is available at https://github.com/SLIT-AI/WRPO. Fanqi Wan, Longguang Zhong, Tianyuan Shi, Xiaojun Quan |
ICLR | 5 |
| 2025 | Discriminative Policy Optimization for Token-Level Reward ModelsabstractProcess reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enhancing the capabilities of LLMs in complex reasoning tasks. Recent efforts have advanced PRMs from step-level to token-level granularity by integrating reward modeling into the training of generative models, with reward scores derived from token generation probabilities. However, the conflict between generative language modeling and reward modeling may introduce instability and lead to inaccurate credit assignments. To address this challenge, we revisit token-level reward assignment by decoupling reward modeling from language generation and derive a token-level reward model through the optimization of a discriminative policy, termed the Q-function Reward Model (Q-RM). We theoretically demonstrate that Q-RM explicitly learns token-level Q-functions from preference data without relying on fine-grained annotations. In our experiments, Q-RM consistently outperforms all baseline methods across various benchmarks.
For example, when integrated into PPO/REINFORCE algorithms, Q-RM enhances the average Pass@1 score by 5.85/4.70 points on mathematical reasoning tasks compared to the ORM baseline, and by 4.56/5.73 points compared to the token-level PRM counterpart. Moreover, reinforcement learning with Q-RM significantly enhances training efficiency, achieving convergence 12× faster than ORM on GSM8K and 11× faster than step-level PRM on MATH. Code and data are available at https://github.com/homzer/Q-RM. Hongzhan Chen, Tao Yang 0033, Shiping Gao, Ruijun Chen 0001, Xiaojun Quan, Hongtao Tian |
ICML | 5 |
| 2025 | Lookahead Routing for Large Language ModelsabstractLarge language model (LLM) routers improve the efficiency of multi-model systems by directing each query to the most appropriate model while leveraging the diverse strengths of heterogeneous LLMs. Most existing approaches frame routing as a classification problem based solely on the input query. While this reduces overhead by avoiding inference across all models, it overlooks valuable information that could be gleaned from potential outputs and fails to capture implicit intent or contextual nuances that often emerge only during response generation. These limitations can result in suboptimal routing decisions, particularly for complex or ambiguous queries that require deeper semantic understanding. To address this challenge, we propose Lookahead, a routing framework that "foresees" potential model outputs by predicting their latent representations and uses these predictions to guide model selection, thus enabling more informed routing without full inference. Within this framework, we implement two approaches based on causal and masked language models. Empirical evaluations across seven public benchmarks—spanning instruction following, mathematical reasoning, and code generation—show that Lookahead consistently outperforms existing routing baselines, achieving an average performance gain of 7.7\% over the state-of-the-art. Our code is available at https://github.com/huangcb01/lookahead-routing. Canbin Huang, Tianyuan Shi, Yuhua Zhu, Ruijun Chen 0001, Xiaojun Quan |
NeurIPS | 5 |
| 2025 | Probabilistic Token Alignment for Large Language Model FusionabstractTraining large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more cost-effective alternative is to fuse existing pre-trained LLMs with different architectures into a more powerful model. However, a key challenge in existing model fusion is their dependence on manually predefined vocabulary alignment, which may not generalize well across diverse contexts, leading to performance degradation in several evaluation. To solve this, we draw inspiration from distribution learning and propose the probabilistic token alignment method as a general and soft mapping for alignment, named as PTA-LLM. Our approach innovatively reformulates token alignment into a classic mathematical problem: optimal transport, seamlessly leveraging distribution-aware learning to facilitate more coherent model fusion. Apart from its inherent generality, PTA-LLM exhibits interpretability from a distributional perspective, offering insights into the essence of the token alignment. Empirical results demonstrate that probabilistic token alignment enhances the target model's performance across multiple capabilities. Runjia Zeng, James Liang, Cheng Han 0001, Zhiwen Cao, Xiaojun Quan, Victor Y. Chen, Lifu Huang, Tong Geng, Qifan Wang 0001, Dongfang Liu |
NeurIPS | 6 |
| 2024 | Small LLMs Are Weak Tool Learners: A Multi-LLM AgentabstractLarge Language Model (LLM) agents significantly extend the capabilities of standalone LLMs, empowering them to interact with external tools (e.g., APIs, functions) and complete various tasks in a self-directed fashion.The challenge of tool use demands that LLMs not only understand user queries and generate answers accurately but also excel in task planning, tool invocation, and result summarization.While traditional works focus on training a single LLM with all these capabilities, performance limitations become apparent, particularly with smaller models.To overcome these challenges, we propose a novel approach that decomposes the aforementioned capabilities into a planner, caller, and summarizer.Each component is implemented by a single LLM that focuses on a specific capability and collaborates with others to accomplish the task.This modular framework facilitates individual updates and the potential use of smaller LLMs for building each capability.To effectively train this framework, we introduce a two-stage training paradigm.First, we fine-tune a backbone LLM on the entire dataset without discriminating sub-tasks, providing the model with a comprehensive understanding of the task.Second, the fine-tuned LLM is used to instantiate the planner, caller, and summarizer respectively, which are continually fine-tuned on respective sub-tasks.Evaluation across various tool-use benchmarks illustrates that our proposed multi-LLM framework surpasses the traditional single-LLM approach, highlighting its efficacy and advantages in tool learning. Weizhou Shen, Chenliang Li 0003, Hongzhan Chen, Ming Yan 0008, Xiaojun Quan, Hehong Chen, Ji Zhang 0011, Fei Huang 0002 |
EMNLP | 5 |
| 2024 | Knowledge Verification to Nip Hallucination in the BudabstractWhile large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.In this paper, we demonstrate the feasibility of mitigating hallucinations by verifying and minimizing the inconsistency between external knowledge present in the alignment data and the intrinsic knowledge embedded within foundation LLMs.Specifically, we propose a novel approach called Knowledge Consistent Alignment (KCA), which employs a well-aligned LLM to automatically formulate assessments based on external knowledge to evaluate the knowledge boundaries of foundation LLMs.To address knowledge inconsistencies in the alignment data, KCA implements several specific strategies to deal with these data instances.We demonstrate the superior efficacy of KCA in reducing hallucinations across six benchmarks, utilizing foundation LLMs of varying backbones and scales.This confirms the effectiveness of mitigating hallucinations by reducing knowledge inconsistency.Our code, model weights, and data are openly accessible at https://github.com/fanqiwan/KCA.* Part of the work was done during his internship at Tencent AI Lab. Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
EMNLP | 4 |
| 2024 | Knowledge Fusion of Large Language ModelsabstractWhile training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures—Llama-2, MPT, and OpenLLaMA—across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at \url{https://github.com/fanqiwan/FuseLLM}. Fanqi Wan, Xinting Huang, Deng Cai 0002, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
ICLR | 4 |
| 2024 | Multi-Party Conversation Modeling for Emotion RecognitionabstractMulti-party conversation modeling plays a vital role in emotion recognition in conversation (ERC). Aside from the intra- and inter-speaker dependencies between different speakers, the difficulty also lies in the fact that each conversation may contain several to many utterances that compose a long text sequence. In this article, we present two approaches to effective multi-party conversation modeling. First, to encode long sequences and capture long-range dependency between utterances, we introduce a dialog-oriented language model, DialogXL, with enhanced memory to store longer conversation sequences and dialog-aware self-attention to deal with multi-party dependencies. Second, we present a directed acyclic neural network, namely DAG-ERC, to encode the utterances with a directed acyclic graph (DAG) to better capture the intrinsic structure within a conversation. DAG-ERC combines the advantages of recurrent models and graph models and provides a more intuitive way to model information flow between sequential utterances. Extensive experiments are conducted on four ERC benchmarks with state-of-the-art models employed for comparison, and empirical results demonstrate the superiority of the two models in multi-party conversation modeling. Xiaojun Quan, Siyue Wu, Weizhou Shen, Jianxing Yu |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Multi-Task Multi-Attention Transformer for Generative Named Entity RecognitionabstractMost previous sequential labeling models are task-specific, while recent years have witnessed the rise of generative models due to the advantage of unifying all named entity recognition (NER) tasks into the encoder-decoder framework. Although achieving promising performance, our pilot studies demonstrate that existing generative models are ineffective at detecting entity boundaries and estimating entity types. In this paper, we propose a multi-task Transformer, which incorporates an entity boundary detection task into the named entity recognition task. More concretely, we achieve entity boundary detection by classifying the relations between tokens within the sentence. To improve the accuracy of entity-type mapping during decoding, we adopt an external knowledge base to calculate the prior entity-type distributions and then incorporate the information into the model via the self- and cross-attention mechanisms. We perform experiments on extensive NER benchmarks, including flat, nested, and discontinuous NER datasets involving long entities. It substantially increases nearly$+0.3 \sim +1.5\;{F_1}$scores across a broad spectrum or performs closely to the best generative NER model. Experimental results show that our approach improves the performance of the generative NER model considerably. Ying Mo, Hongyin Tang, Qifan Wang 0001, Zenglin Xu, Jingang Wang, Xiaojun Quan, Wei Wu 0014, Zhoujun Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | A Graph Fusion Approach for Cross-Lingual Machine Reading ComprehensionabstractAlthough great progress has been made for Machine Reading Comprehension (MRC) in English, scaling out to a large number of languages remains a huge challenge due to the lack of large amounts of annotated training data in non-English languages. To address this challenge, some recent efforts of cross-lingual MRC employ machine translation to transfer knowledge from English to other languages, through either explicit alignment or implicit attention. For effective knowledge transition, it is beneficial to leverage both semantic and syntactic information. However, the existing methods fail to explicitly incorporate syntax information in model learning. Consequently, the models are not robust to errors in alignment and noises in attention. In this work, we propose a novel approach, which jointly models the cross-lingual alignment information and the mono-lingual syntax information using a graph. We develop a series of algorithms, including graph construction, learning, and pre-training. The experiments on two benchmark datasets for cross-lingual MRC show that our approach outperforms all strong baselines, which verifies the effectiveness of syntax information for cross-lingual MRC. Zenan Xu, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Qinliang Su, Xiaojun Quan, Daxin Jiang |
AAAI | 6 |
| 2023 | Orders Are Unwanted: Dynamic Deep Graph Convolutional Network for Personality DetectionabstractPredicting personality traits based on online posts has emerged as an important task in many fields such as social network analysis. One of the challenges of this task is assembling information from various posts into an overall profile for each user. While many previous solutions simply concatenate the posts into a long text and then encode the text by sequential or hierarchical models, they introduce unwarranted orders for the posts, which may mislead the models. In this paper, we propose a dynamic deep graph convolutional network (D-DGCN) to overcome the above limitation. Specifically, we design a learn-to-connect approach that adopts a dynamic multi-hop structure instead of a deterministic structure, and combine it with the DGCN module to automatically learn the connections between posts. The modules of post encoder, learn-to-connect, and DGCN are jointly trained in an end-to-end manner. Experimental results on the Kaggle and Pandora datasets show the superior performance of D-DGCN to state-of-the-art baselines. Our code is available at https://github.com/djz233/D-DGCN. Tao Yang 0033, Jinghao Deng, Xiaojun Quan, Qifan Wang 0001 |
AAAI | 3 |
| 2023 | Disentangled Phonetic Representation for Chinese Spelling CorrectionabstractChinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts.Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phonetic representations with character representations, which tends to weaken the representation effect of normal texts.In this work, we propose to disentangle the two types of features to allow for direct interaction between textual and phonetic information.To learn useful phonetic representations, we introduce a pinyin-to-character objective to ask the model to predict the correct characters based solely on phonetic information, where a separation mask is imposed to disable attention from phonetic input to text.To avoid overfitting the phonetics, we further design a self-distillation module to ensure that semantic information plays a major role in the prediction.Extensive experiments on three CSC benchmarks demonstrate the superiority of our method in using phonetic information 1 . Zihong Liang, Xiaojun Quan, Qifan Wang 0001 |
ACL (1) | 2 |
| 2023 | Multi-Grained Knowledge Retrieval for End-to-End Task-Oriented DialogabstractRetrieving proper domain knowledge from an external database lies at the heart of end-toend task-oriented dialog systems to generate informative responses.Most existing systems blend knowledge retrieval with response generation and optimize them with direct supervision from reference responses, leading to suboptimal retrieval performance when the knowledge base becomes large-scale.To address this, we propose to decouple knowledge retrieval from response generation and introduce a multigrained knowledge retriever (MAKER) that includes an entity selector to search for relevant entities and an attribute selector to filter out irrelevant attributes.To train the retriever, we propose a novel distillation objective that derives supervision signals from the response generator.Experiments conducted on three standard benchmarks with both small and largescale knowledge bases demonstrate that our retriever performs knowledge retrieval more effectively than existing methods.Our code has been made publicly available. Fanqi Wan, Weizhou Shen, Xiaojun Quan, Wei Bi |
ACL (1) | 4 |
| 2023 | MUSTIE: Multimodal Structural Transformer for Web Information ExtractionabstractQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qifan Wang 0001, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu |
ACL (1) | 3 |
| 2023 | AD-KD: Attribution-Driven Knowledge Distillation for Language Model CompressionabstractKnowledge distillation has attracted a great deal of interest recently to compress pre-trained language models.However, existing knowledge distillation methods suffer from two limitations.First, the student model simply imitates the teacher's behavior while ignoring the underlying reasoning.Second, these methods usually focus on the transfer of sophisticated model-specific knowledge but overlook dataspecific knowledge.In this paper, we present a novel attribution-driven knowledge distillation approach, which explores the token-level rationale behind the teacher model based on Integrated Gradients (IG) and transfers attribution knowledge to the student model.To enhance the knowledge transfer of model reasoning and generalization, we further explore multi-view attribution distillation on all potential decisions of the teacher.Comprehensive experiments are conducted with BERT on the GLUE benchmark.The experimental results demonstrate the superior performance of our approach to several state-of-the-art methods. Siyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 0001, Rui Wang 0005 |
ACL (1) | 3 |
| 2023 | Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue SystemabstractDeveloping an efficient retriever to retrieve knowledge from a large-scale knowledge base (KB) is critical for task-oriented dialogue systems to effectively handle localized and specialized tasks.However, widely used generative models such as T5 and ChatGPT often struggle to differentiate subtle differences among the retrieved KB records when generating responses, resulting in suboptimal quality of generated responses.In this paper, we propose the application of maximal marginal likelihood to train a perceptive retriever by utilizing signals from response generation for supervision.In addition, our approach goes beyond considering solely retrieved entities and incorporates various meta knowledge to guide the generator, thus improving the utilization of knowledge.We evaluate our approach on three task-oriented dialogue datasets using T5 and ChatGPT as the backbone models.The results demonstrate that when combined with meta knowledge, the response generator can effectively leverage high-quality knowledge records from the retriever and enhance the quality of generated responses.The code of this work is available at https://github.com/shenwzh3/MK-TOD. Weizhou Shen, Yingqi Gao, Canbin Huang, Fanqi Wan, Xiaojun Quan, Wei Bi |
EMNLP | 5 |
| 2023 | Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue SystemsabstractEfficient knowledge retrieval plays a pivotal role in ensuring the success of end-to-end taskoriented dialogue systems by facilitating the selection of relevant information necessary to fulfill user requests.However, current approaches generally integrate knowledge retrieval and response generation, which poses scalability challenges when dealing with extensive knowledge bases.Taking inspiration from open-domain question answering, we propose a retrievergenerator architecture that harnesses a retriever to retrieve pertinent knowledge and a generator to generate system responses.Due to the lack of retriever training labels, we propose relying on feedback from the generator as pseudo-labels to train the retriever.To achieve this, we introduce a dual-feedback mechanism that generates both positive and negative feedback based on the output of the generator.Our method demonstrates superior performance in task-oriented dialogue tasks, as evidenced by experimental results on three benchmark datasets. Tianyuan Shi, Liangzhi Li 0004, Zijian Lin, Tao Yang 0033, Xiaojun Quan, Qifan Wang 0001 |
EMNLP | 5 |
| 2023 | Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active ExplorationabstractInstruction-tuning can be substantially optimized through enhanced diversity, resulting in models capable of handling a broader spectrum of tasks.However, existing data employed for such tuning often exhibit an inadequate coverage of individual domains, limiting the scope for nuanced comprehension and interactions within these areas.To address this deficiency, we propose EXPLORE-INSTRUCT, a novel approach to enhance the data coverage to be used in domain-specific instruction-tuning through active exploration via Large Language Models (LLMs).Built upon representative domain use cases, EXPLORE-INSTRUCT explores a multitude of variations or possibilities by implementing a search algorithm to obtain diversified and domain-focused instruction-tuning data.Our data-centric analysis validates the effectiveness of this proposed approach in improving domain-specific instruction coverage.Moreover, our model's performance demonstrates considerable advancements over multiple baselines, including those utilizing domainspecific data enhancement.Our findings offer a promising opportunity to improve instruction coverage, especially in domain-specific contexts, thereby advancing the development of adaptable language models.Our code, model weights, and data are public at https:// github.com/fanqiwan/Explore-Instruct. Fanqi Wan, Xinting Huang, Tao Yang 0033, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
EMNLP | 4 |
| 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsabstractQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu |
EMNLP | 9 |
| 2023 | Generic Dependency Modeling for Multi-Party ConversationabstractTo model the dependencies between utterances in multi-party conversations, we propose a simple and generic framework based on the dependency parsing results of utterances. Particularly, we present an approach to encoding the dependencies in the form of relative dependency encoding (ReDE) and illustrate how to implement it in Transformers by modifying the computation of self-attention. Experimental results on four multi-party conversation benchmarks show that this framework successfully boosts the general performance of two Transformer-based language models and leads to comparable or even superior performance compared to the state-of-the-art methods. The code is available at https://github.com/shenwzh3/ReDE. Weizhou Shen, Xiaojun Quan |
ICASSP | 2 |
| 2023 | A Unified Generation Approach for Robust Dialogue State Tracking
Zijian Lin, Beizhang Guo, Tianyuan Shi, Xiaojun Quan, Liangzhi Li 0004 |
NLPCC (1) | 5 |
| 2023 | Compound Aspect Extraction by Augmentation and Constituency LatticeabstractAspects are opinion targets to extract in aspect-based sentiment analysis. While existing methods can already produce satisfactory extraction results, they suffer when faced with compound aspect terms, typically phrase-level aspect terms that have inner structure and occur infrequently in the training set. This issue can be mainly attributed to the scarcity of training examples targeting compound aspect terms and by the neglect of the syntactic structure of a sentence in the modeling process. In this article, we aim to cope with compound aspect extraction by a two-stage hybrid approach. First, we introduce a conditional generation method for data augmentation in a masked sequence-to-sequence framework, which is controllable to preserve original aspects while generating a new sentence. Second, we propose a constituency lattice structure that is induced from the constituency-based parse tree of a sentence. Experimental results on two review datasets show that this approach can greatly improve the effect of compound aspect extraction. Xiaojun Quan, Zhengcheng Min, Kun Li 0003, Yunyi Yang |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Multi-Hop Reasoning Question Generation and Its ApplicationabstractThis article focuses on the topic of multi-hop question generation (QG), which aims to generate the questions requiring multi-hop reasoning skills from the given text. These questions are not only syntactically valid but also logically correlated with the answers. Concretely, we first design a basic QG model and customize several techniques to ensure results' syntactic validity. In order to promote the logical correlations, we use a reasoning chain extracted from the text to regularize the results. Considering that different samples have their own characteristics on the aspects of text contextual structure, the type of question, and logical correlation, we propose a new adaptive meta-learner to optimize the basic QG model. Each case and its similar samples are viewed as a pseudo-QG task. The similar structural contexts contained in the same task are used as guidance to fine-tune the model. To measure the similarity of samples' structured inputs, we propose a data-driven multi-level recognizer. The experimental results on two typical data sets in various domains show the effectiveness of the proposed approach. Moreover, we apply the generated results to the task of machine reading comprehension and achieve significant performance improvements. That demonstrates the capacity of multi-hop QG in facilitating real-world applications. Jianxing Yu, Qinliang Su, Xiaojun Quan, Jian Yin 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Autoregressive Entity Generation for End-to-End Task-Oriented DialogabstractTask-oriented dialog (TOD) systems are often required to interact with an external knowledge base (KB) to retrieve necessary entity (e.g., restaurants) information to support their response generation. Most current end-to-end TOD systems either retrieve the KB information explicitly or embed it into model parameters for implicit access. While the first approach demands scanning the KB at each turn of response generation, which is inefficient when the KB scales up, the second approach shows higher flexibility and efficiency. In either approach, the response shall contain attributes of the same entity, however the systems may generate a response with conflicting entities. To address this, we propose to generate the entity autoregressively before leveraging it to guide the response generation in an end-to-end system. To ensure entity consistency, we impose a trie constraint on the decoding of an entity. We also introduce a logit concatenation strategy to facilitate gradient backpropagation for end-to-end training. Experiments on MultiWOZ 2.1 single and CAMREST show that our system can generate more high-quality and entity-consistent responses in an end-to-end manner. Guanhuan Huang, Xiaojun Quan, Qifan Wang 0001 |
COLING | 2 |
| 2022 | XPrompt: Exploring the Extreme of Prompt TuningabstractPrompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner.While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performance gap between prompt tuning and fine-tuning for models of moderate and small scales (typically less than 11B parameters).In this paper, we empirically show that the trained prompt tokens can have a negative impact on a downstream task and thus degrade its performance.To bridge the gap, we propose a novel PROMPT tuning model with an eXtremely small scale (XPROMPT) under the regime of lottery tickets hypothesis.Specifically, XPROMPT eliminates the negative prompt tokens at different granularity levels through a hierarchical structured pruning, yielding a more parameter-efficient prompt yet with a competitive performance.Comprehensive experiments are carried out on the SuperGLUE tasks, and the results indicate that XPROMPT is able to close the performance gap at smaller model scales. 1 Recently, Prompt-Tuning (Lester et al., 2021; Liu et al., 2021b) has been proposed to address this issue by prepending a soft prompt to the input and only updating the parameters of prompt tokens during tuning.Prompt-Tuning provides a parameter-efficient alternative to fine-tuning, since the scale of the soft prompt is tens of thousand smaller.It is also conceptually simpler and more flexible than other parameter-efficient tuning methods (such as Adapters), that require intrusive modifications to transformer layers (Houlsby et al., 2019;Guo et al., 2021).Using fewer tunable parameters, prompt tuning achieves competitive performance to fine-tuning with the increase of the model scale.However, there is still a large performance gap between prompt tuning and fine-tuning for models of smaller scales (as shown in Figure 1).This paper aims to fill the gap, from the perspective of the lottery tickets hypothesis (LTH) (Frankle and Carbin, 2019).We are motivated by an observation that, on a specific task, not all prompt tokens contribute equally to the task performance, while certain prompt tokens may even bring a negative Fang Ma, Chen Zhang 0020, Jingang Wang, Qifan Wang 0001, Wei Wu 0014, Xiaojun Quan, Dawei Song 0001 |
EMNLP | 7 |
| 2022 | Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word GenerationabstractAutomatic question generation (AQG) is the task of generating a question from a given passage and an answer.Most existing AQG methods aim at encoding the passage and the answer to generate the question.However, limited work has focused on modeling the correlation between the target answer and the generated question.Moreover, unseen or rare word generation has not been studied in previous works.In this paper, we propose a novel approach which incorporates question generation with its dual problem, question answering, into a unified primal-dual framework.Specifically, the question generation component consists of an encoder that jointly encodes the answer with the passage, and a decoder that produces the question.The question answering component then re-asks the generated question on the passage to ensure that the target answer is obtained.We further introduce a knowledge distillation module to improve the model generalization ability.We conduct an extensive set of experiments on SQuAD and HotpotQA benchmarks.Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods. Qifan Wang 0001, Xiaojun Quan, Fuli Feng, Dongfang Liu, Zenglin Xu, Sinong Wang, Hao Ma 0001 |
EMNLP | 3 |
| 2022 | GL-RG: Global-Local Representation Granularity for Video CaptioningabstractVideo captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local representation across video frames for caption generation, leaving plenty of room for improvement. In this work, we approach the video captioning task from a new perspective and propose a GL-RG framework for video captioning, namely a Global-Local Representation Granularity. Our GL-RG demonstrates three advantages over the prior efforts: 1) we explicitly exploit extensive visual representations from different video ranges to improve linguistic expression; 2) we devise a novel global-local encoder to produce rich semantic vocabulary to obtain a descriptive granularity of video contents across frames; 3) we develop an incremental training strategy which organizes model learning in an incremental fashion to incur an optimal captioning behavior. Experimental results on the challenging MSR-VTT and MSVD datasets show that our DL-RG outperforms recent state-of-the-art methods by a significant margin. Code is available at https://github.com/ylqi/GL-RG. Liqi Yan, Qifan Wang 0001, Yiming Cui 0002, Fuli Feng, Xiaojun Quan, Xiangyu Zhang 0001, Dongfang Liu |
IJCAI | 5 |
| 2022 | AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningabstractFine-tuning large pre-trained language models on downstream tasks is apt to suffer from overfitting when limited training data is available. While dropout proves to be an effective antidote by randomly dropping a proportion of units, existing research has not examined its effect on the self-attention mechanism. In this paper, we investigate this problem through self-attention attribution and find that dropping attention positions with low attribution scores can accelerate training and increase the risk of overfitting. Motivated by this observation, we propose Attribution-Driven Dropout (AD-DROP), which randomly discards some high-attribution positions to encourage the model to make predictions by relying more on low-attribution positions to reduce overfitting. We also develop a cross-tuning strategy to alternate fine-tuning and AD-DROP to avoid dropping high-attribution positions excessively. Extensive experiments on various benchmarks show that AD-DROP yields consistent improvements over baselines. Analysis further confirms that AD-DROP serves as a strategic regularizer to prevent overfitting during fine-tuning. Tao Yang 0033, Jinghao Deng, Xiaojun Quan, Qifan Wang 0001, Shaoliang Nie |
NeurIPS | 3 |
| 2022 | WebFormer: The Web-page Transformer for Structure Information ExtractionabstractStructure information extraction refers to the task of extracting structured text fields from web pages, such as extracting a product offer from a shopping page including product title, description, brand and price. It is an important research topic which has been widely studied in document understanding and web search. Recent natural language models with sequence modeling have demonstrated state-of-the-art performance on web information extraction. However, effectively serializing tokens from unstructured web pages is challenging in practice due to a variety of web layout patterns. Limited work has focused on modeling the web layout for extracting the text fields. In this paper, we introduce WebFormer, a Web-page transFormer model for structure information extraction from web documents. First, we design HTML tokens for each DOM node in the HTML by embedding representations from their neighboring tokens through graph attention. Second, we construct rich attention patterns between HTML tokens and text tokens, which leverages the web layout for effective attention weight computation. We conduct an extensive set of experiments on SWDE and Common Crawl benchmarks. Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods. Qifan Wang 0001, Yi Fang 0008, Anirudh Ravula, Fuli Feng, Xiaojun Quan, Dongfang Liu |
WWW | 5 |
| 2021 | DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion RecognitionabstractThis paper presents our pioneering effort for emotion recognition in conversation (ERC) with pre-trained language models. Unlike regular documents, conversational utterances appear alternately from different parties and are usually organized as hierarchical structures in previous work. Such structures are not conducive to the application of pre-trained language models such as XLNet. To address this issue, we propose an all-in-one XLNet model, namely DialogXL, with enhanced memory to store longer historical context and dialog-aware self-attention to deal with the multi-party structures. Specifically, we first modify the recurrence mechanism of XLNet from segment-level to utterance-level in order to better model the conversational data. Second, we introduce dialog-aware self-attention in replacement of the vanilla self-attention in XLNet to capture useful intra- and inter-speaker dependencies. Extensive experiments are conducted on four ERC benchmarks with mainstream models presented for comparison. The experimental results show that the proposed model outperforms the baselines on all the datasets. Several other experiments such as ablation study and error analysis are also conducted and the results confirm the role of the critical modules of DialogXL. Weizhou Shen, Xiaojun Quan, Zhixian Xie |
AAAI | 3 |
| 2021 | UBAR: Towards Fully End-to-End Task-Oriented Dialog System with GPT-2abstractThis paper presents our task-oriented dialog system UBAR which models task-oriented dialogs on a dialog session level. Specifically, UBAR is acquired by fine-tuning the large pre-trained unidirectional language model GPT-2 on the sequence of the entire dialog session which is composed of user utterance, belief state, database result, system act, and system response of every dialog turn. Additionally, UBAR is evaluated in a more realistic setting, where its dialog context has access to user utterances and all content it generated such as belief states, system acts, and system responses. Experimental results on the MultiWOZ datasets show that UBAR achieves state-of-the-art performances in multiple settings, improving the combined score of response generation, policy optimization, and end-to-end modeling by 4.7, 3.5, and 9.4 points respectively. Thorough analyses demonstrate that the session-level training sequence formulation and the generated dialog context are essential for UBAR to operate as a fully end-to-end task-oriented dialog system in real life. We also examine the transfer ability of UBAR to new domains with limited data and provide visualization and a case study to illustrate the advantages of UBAR in modeling on a dialog session level. Yunyi Yang, Xiaojun Quan |
AAAI | 3 |
| 2021 | Multi-Document Transformer for Personality DetectionabstractPersonality detection aims to identify the personality traits implied in social media posts. The core of this task is to put together information in multiple scattered posts to depict an overall personality profile for each user. Existing approaches either encode each post individually or assemble posts arbitrarily into a new document that can be encoded sequentially or hierarchically. While the first approach ignores the connection between posts, the second tends to introduce unnecessary post-order bias into posts. In this paper, we propose a multi-document Transformer, namely Transformer-MD, to tackle the above issues. When encoding each post, Transformer-MD allows access to information in the other posts of the user through Transformer-XL’s memory tokens which share the same position embedding.Besides, personality is usually defined along different traits and each trait may need to attend to different post information, which has rarely been touched by existing research. To address this concern, we propose a dimension attention mechanism on top of Transformer-MD to obtain trait-specific representations for multi-trait personality detection. We evaluate the proposed model on the Kaggle and Pandora MBTI datasets and the experimental results show that it compares favorably with baseline methods. Xiaojun Quan, Yunyi Yang, Jianxing Yu |
AAAI | 2 |
| 2021 | Directed Acyclic Graph Network for Conversational Emotion RecognitionabstractWeizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Weizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan |
ACL/IJCNLP (1) | 4 |
| 2021 | Syntax-Enhanced Pre-trained ModelabstractZenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong 0001, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan 0001 |
ACL/IJCNLP (1) | 8 |
| 2021 | Psycholinguistic Tripartite Graph Network for Personality DetectionabstractTao Yang, Feifan Yang, Haolan Ouyang, Xiaojun Quan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tao Yang 0033, Haolan Ouyang, Xiaojun Quan |
ACL/IJCNLP (1) | 4 |
| 2021 | Progressive Dialogue State Tracking for Multi-Domain Dialogue Systems
Minqian Liu, Xiaojun Quan |
ICASSP | 3 |
| 2020 | Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence GenerationabstractAspect term extraction aims to extract aspect terms from review texts as opinion targets for sentiment analysis.One of the big challenges with this task is the lack of sufficient annotated data.While data augmentation is potentially an effective technique to address the above issue, it is uncontrollable as it may change aspect words and aspect labels unexpectedly.In this paper, we formulate the data augmentation as a conditional generation task: generating a new sentence while preserving the original opinion targets and labels.We propose a masked sequence-to-sequence method for conditional augmentation of aspect term extraction.Unlike existing augmentation approaches, ours is controllable and allows us to generate more diversified sentences.Experimental results confirm that our method alleviates the data scarcity problem significantly.It also effectively boosts the performances of several current models for aspect term extraction. Kun Li 0003, Chengbo Chen, Xiaojun Quan, Yan Song 0003 |
ACL | 3 |
| 2020 | Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed KnowledgeabstractChinese word segmentation (CWS) and partof-speech (POS) tagging are important fundamental tasks for Chinese language processing, where joint learning of them is an effective one-step solution for both tasks.Previous studies for joint CWS and POS tagging mainly follow the character-based tagging paradigm with introducing contextual information such as n-gram features or sentential representations from recurrent neural models.However, for many cases, the joint tagging needs not only modeling from context features but also knowledge attached to them (e.g., syntactic relations among words); limited efforts have been made by existing research to meet such needs.In this paper, we propose a neural model named TWASP for joint CWS and POS tagging following the character-based sequence labeling paradigm, where a two-way attention mechanism is used to incorporate both context feature and their corresponding syntactic knowledge for each input character.Particularly, we use existing language processing toolkits to obtain the auto-analyzed syntactic knowledge for the context, and the proposed attention module can learn and benefit from them although their quality may not be perfect.Our experiments illustrate the effectiveness of the two-way attentions for joint CWS and POS tagging, where state-of-the-art performance is achieved on five benchmark datasets.1 Yuanhe Tian, Yan Song 0003, Xiang Ao 0001, Fei Xia 0004, Xiaojun Quan, Tong Zhang 0001 |
ACL | 5 |
| 2020 | Relational Graph Attention Network for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis aims to determine the sentiment polarity towards a specific aspect in online reviews.Most recent efforts adopt attention-based neural network models to implicitly connect aspects with opinion words.However, due to the complexity of language and the existence of multiple aspects in a single sentence, these models often confuse the connections.In this paper, we address this problem by means of effective encoding of syntax information.Firstly, we define a unified aspect-oriented dependency tree structure rooted at a target aspect by reshaping and pruning an ordinary dependency parse tree.Then, we propose a relational graph attention network (R-GAT) to encode the new tree structure for sentiment prediction.Extensive experiments are conducted on the SemEval 2014 and Twitter datasets, and the experimental results confirm that the connections between aspects and opinion words can be better established with our approach, and the performance of the graph attention network (GAT) is significantly improved as a consequence. Weizhou Shen, Yunyi Yang, Xiaojun Quan, Rui Wang 0005 |
ACL | 4 |
| 2020 | Multi-Domain Dialogue Acts and Response Co-GenerationabstractGenerating fluent and informative responses is of critical importance for task-oriented dialogue systems.Existing pipeline approaches generally predict multiple dialogue acts first and use them to assist response generation.There are at least two shortcomings with such approaches.First, the inherent structures of multi-domain dialogue acts are neglected.Second, the semantic associations between acts and responses are not taken into account for response generation.To address these issues, we propose a neural co-generation model that generates dialogue acts and responses concurrently.Unlike those pipeline approaches, our act generation module preserves the semantic structures of multi-domain dialogue acts and our response generation module dynamically attends to different acts as needed.We train the two modules jointly using an uncertainty loss to adjust their task weights adaptively.Extensive experiments are conducted on the largescale MultiWOZ dataset and the results show that our model achieves very favorable improvement over several state-of-the-art models in both automatic and human evaluations. Rui Wang 0005, Xiaojun Quan, Jianxing Yu |
ACL | 4 |
| 2020 | Low-Resource Generation of Multi-hop Reasoning QuestionsabstractThis paper focuses on generating multi-hop reasoning questions from the raw text in a low resource circumstance.Such questions have to be syntactically valid and need to logically correlate with the answers by deducing over multiple relations on several sentences in the text.Specifically, we first build a multi-hop generation model and guide it to satisfy the logical rationality by the reasoning chain extracted from a given text.Since the labeled data is limited and insufficient for training, we propose to learn the model with the help of a large scale of unlabeled data that is much easier to obtain.Such data contains rich expressive forms of the questions with structural patterns on syntax and semantics.These patterns can be estimated by the neural hidden semi-Markov model using latent variables.With latent patterns as a prior, we can regularize the generation model and produce the optimal results.Experimental results on the HotpotQA data set demonstrate the effectiveness of our model.Moreover, we apply the generated results to the task of machine reading comprehension and achieve significant performance improvements. Jianxing Yu, Wei Liu 0061, Qinliang Su, Xiaojun Quan, Jian Yin 0001 |
ACL | 6 |
| 2020 | Multi-choice Relational Reasoning for Machine Reading ComprehensionabstractThis paper presents our study of cloze-style reading comprehension by imitating human reading comprehension, which normally involves tactical comparing and reasoning over candidates while choosing the best answer. We propose a multi-choice relational reasoning (McR2) model with an aim to enable relational reasoning on candidates based on fusion representations of document, query and candidates. For the fusion representations, we develop an efficient encoding architecture by integrating the schemes of bidirectional attention flow, self-attention and document-gated query reading. Then, comparing and inferring over candidates are executed by a novel relational reasoning network. We conduct extensive experiments on four datasets derived from two public corpora, Children’s Book Test and Who DiD What, to verify the validity and advantages of our model. The results show that it outperforms all baseline models significantly on the four benchmark datasets. The effectiveness of its key components is also validated by an ablation study. Wuya Chen, Xiaojun Quan, Chunyu Kit, Zhengcheng Min, Jiahai Wang |
COLING | 2 |
| 2020 | Embedding Dynamic Attributed Networks by Modeling the Evolution ProcessesabstractNetwork embedding has recently emerged as a promising technique to embed nodes of a network into low-dimensional vectors.While fairly successful, most existing works focus on the embedding techniques for static networks.But in practice, there are many networks that are evolving over time and hence are dynamic, e.g., the social networks.To address this issue, a high-order spatio-temporal embedding model is developed to track the evolutions of dynamic networks.Specifically, an activeness-aware neighborhood embedding method is first proposed to extract the high-order neighborhood information at each given timestamp.Then, an embedding prediction framework is further developed to capture the temporal correlations, in which the attention mechanism is employed instead of recurrent neural networks (RNNs) for its efficiency in computing and flexibility in modeling.Extensive experiments are conducted on four realworld datasets from three different areas.It is shown that the proposed method outperforms all the baselines by a substantial margin for the tasks of dynamic link prediction and node classification, which demonstrates the effectiveness of the proposed methods on tracking the evolutions of dynamic networks. Zenan Xu, Zijing Ou, Qinliang Su, Jianxing Yu, Xiaojun Quan, Zhenkun Lin |
COLING | 5 |
| 2020 | Constituency Lattice Encoding for Aspect Term ExtractionabstractOne of the remaining challenges for aspect term extraction in sentiment analysis resides in the extraction of phrase-level aspect terms, which is non-trivial to determine the boundaries of such terms. In this paper, we aim to address this issue by incorporating the span annotations of constituents of a sentence to leverage the syntactic information in neural network models. To this end, we first construct a constituency lattice structure based on the constituents of a constituency tree. Then, we present two approaches to encoding the constituency lattice using BiLSTM-CRF and BERT as the base models, respectively. We experimented on two benchmark datasets to evaluate the two models, and the results confirm their superiority with respective 3.17 and 1.35 points gained in F1-Measure over the current state of the art. The improvements justify the effectiveness of the constituency lattice for aspect term extraction. Yunyi Yang, Kun Li 0003, Xiaojun Quan, Weizhou Shen, Qinliang Su |
COLING | 3 |
| 2020 | Generating Multi-hop Reasoning Questions to Improve Machine Reading ComprehensionabstractThis paper focuses on the topic of multi-hop question generation, which aims to generate questions needed reasoning over multiple sentences and relations to derive answers. In particular, we first build an entity graph to integrate various entities scattered over text based on their contextual relations. We then heuristically extract the sub-graph by the evidential relations and type, so as to obtain the reasoning chain and textual related contents for each question. Guided by the chain, we propose a holistic generator-evaluator network to form the questions, where such guidance helps to ensure the rationality of generated questions which need multi-hop deduction to correspond to the answers. The generator is a sequence-to-sequence model, designed with several techniques to make the questions syntactically and semantically valid. The evaluator optimizes the generator network by employing a hybrid mechanism combined of supervised and reinforced learning. Experimental results on HotpotQA data set demonstrate the effectiveness of our approach, where the generated samples can be used as pseudo training data to alleviate the data shortage problem for neural network and assist to learn the state-of-the-arts for multi-hop machine comprehension. Jianxing Yu, Xiaojun Quan, Qinliang Su, Jian Yin 0001 |
WWW | 2 |
| 2019 | BiSET: Bi-directional Selective Encoding with Template for Abstractive SummarizationabstractThe success of neural summarization models stems from the meticulous encodings of source articles.To overcome the impediments of limited and sometimes noisy training data, one promising direction is to make better use of the available training data by applying filters during summarization.In this paper, we propose a novel Bi-directional Selective Encoding with Template (BiSET) model, which leverages template discovered from training data to softly select key information from each source article to guide its summarization process.Extensive experiments on a standard summarization dataset were conducted and the results show that the template-equipped BiSET model manages to improve the summarization performance significantly with a new state of the art. Xiaojun Quan, Rui Wang 0005 |
ACL (1) | 2 |
| 2019 | A Deep Neural Information Fusion Architecture for Textual Network EmbeddingsabstractZenan Xu, Qinliang Su, Xiaojun Quan, Weijia Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zenan Xu, Qinliang Su, Xiaojun Quan |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Gated Convolutional Networks for Commonsense Machine Comprehension
Wuya Chen, Xiaojun Quan, Chengbo Chen |
ICONIP (1) | 2 |
| 2015 | Regularizing Flat Latent Variables with Hierarchical Structures
Rongcheng Lin, Xiaojun Quan, Richang Hong, Zhiang Wu 0001, Yong Ge 0001 |
IJCAI | 3 |
| 2015 | Short and Sparse Text Topic Modeling via Self-Aggregation
Xiaojun Quan, Chunyu Kit, Yong Ge 0001, Sinno Jialin Pan |
IJCAI | 1 |
| 2015 | Towards non-monotonic sentence alignment
Xiaojun Quan, Chunyu Kit |
Inf. Sci. | 1 |
| 2015 | Latent Discriminative Models for Social Emotion Detection with Emotional DependencyabstractSentiment analysis of such opinionated online texts as reviews and comments has received increasingly close attention, yet most of the work is intended to deal with the detection of authors’ emotion. In contrast, this article presents our study of the social emotion detection problem, the objective of which is to identify the evoked emotions of readers by online documents such as news articles. A novel Latent Discriminative Model (LDM) is proposed for this task. LDM works by introducing intermediate hidden variables to model the latent structure of input text corpora. To achieve this, it defines a joint distribution over emotions and latent variables, conditioned on the observed text documents. Moreover, we assume that social emotions are not independent but correlated with one another, and the dependency of them is capable of providing additional guidance to LDM in the training process. The inclusion of this emotional dependency into LDM gives rise to a new Emotional Dependency-based LDM (eLDM). We evaluate the proposed models through a series of empirical evaluations on two real-world corpora of news articles. Experimental results verify the effectiveness of LDM and eLDM in social emotion detection. Xiaojun Quan, Qifan Wang 0001, Ying Zhang 0015, Luo Si, Wenyin Liu |
ACM Trans. Inf. Syst. | 1 |
| 2014 | Towards building a social emotion detection system for online news
Jingsheng Lei, Yanghui Rao, Qing Li 0001, Xiaojun Quan, Wenyin Liu |
Future Gener. Comput. Syst. | 4 |
| 2014 | Affective topic model for social emotion detection
Yanghui Rao, Qing Li 0001, Wenyin Liu, Qingyuan Wu, Xiaojun Quan |
Neural Networks | 5 |
| 2013 | Non-Monotonic Sentence Alignment via Semisupervised Learning
Xiaojun Quan, Chunyu Kit, Yan Song 0003 |
ACL (1) | 1 |
| 2013 | Feature selection for high-dimensional imbalanced data
Liuzhi Yin, Yong Ge 0001, Keli Xiao, Xiaojun Quan |
Neurocomputing | 5 |
| 2012 | Emotion tagging for comments of online news by meta classification with heterogeneous information sourcesabstractWith the rapid growth of online news services, users can actively respond to online news by making comments. Users often express subjective emotions in comments such as sadness, surprise and anger. Such emotions can help understand the preferences and perspectives of individual users, and therefore may facilitate online publishers to provide users with more relevant services. This paper tackles the task of predicting emotions for the comments of online news. To the best of our knowledge, this is the first research work for addressing the task. In particular, this paper proposes a novel Meta classification approach that exploits heterogeneous information sources such as the content of the comments and the emotion tags of news articles generated by users. The experiments on two datasets from online news services demonstrate the effectiveness of the proposed approach. Ying Zhang 0015, Yi Fang 0008, Xiaojun Quan, Luo Si, Xiaojie Yuan |
SIGIR | 3 |
| 2012 | User interest modeling and its application for question recommendation in user-interactive question answering systems
Xingliang Ni, Xiaojun Quan, Wenyin Liu, Bei Hua |
Inf. Process. Manag. | 3 |
| 2011 | Automatic categorization of questions for user-interactive question answering
Wanpeng Song, Wenyin Liu, Naijie Gu, Xiaojun Quan, Tianyong Hao |
Inf. Process. Manag. | 4 |
| 2011 | Short text clustering by finding core terms
Xingliang Ni, Xiaojun Quan, Wenyin Liu, Bei Hua |
Knowl. Inf. Syst. | 2 |
| 2011 | Term Weighting Schemes for Question CategorizationabstractTerm weighting has proven to be an effective way to improve the performance of text categorization. Very recently, with the development of user-interactive question answering or community question answering, there has emerged a need to accurately categorize questions into predefined categories. However, as a question is usually a piece of short text, can the existing term-weighting methods perform consistently in question categorization as they do in text categorization? The answer is not clear, since to the best of our knowledge, we have not seen any work related to this problem despite of its significance. In this study, we investigate the popular unsupervised and supervised term-weighting methods for question categorization. At the same time, we propose three new supervised term-weighting methods, namely, qf*icf, iqf*qf*icf, and vrf. Comparisons of them with existing unsupervised and supervised term-weighting methods are made through a series of experiments on question collections of Yahoo! Answers. The experimental results show that iqf*qf*icf achieves the best performance among all term-weighting methods, while qf*icf and vrf are also competitive for question categorization. Meanwhile, tf*OR is proven to be the most significant one among existing methods. In addition, iqf*qf*icf and vrf are also effective for long document categorization. Xiaojun Quan, Wenyin Liu, Bite Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Exploring the Sentiment Strength of User Reviews
Xiangfei Kong, Xiaojun Quan, Wenyin Liu, Yinlong Xu 0001 |
WAIM | 3 |
| 2010 | Discovering phishing target based on semantic link network
Wenyin Liu, Xiaojun Quan, Bite Qiu, Gang Liu 0008 |
Future Gener. Comput. Syst. | 3 |
| 2010 | A short text modeling method combining semantic and statistical information
Wenyin Liu, Xiaojun Quan, Bite Qiu |
Inf. Sci. | 2 |
| 2010 | Short text similarity based on probabilistic topics
Xiaojun Quan, Gang Liu 0008, Xingliang Ni, Wenyin Liu |
Knowl. Inf. Syst. | 1 |
| 2008 | Adaptive label-driven scaling for latent semantic indexingabstractThis paper targets on enhancing Latent Semantic Indexing (LSI) by exploiting category labels. Specifically, in the term-document matrix, the vector for each term either appearing in labels or semantically close to labels is scaled before performing Singular Value Decomposition (SVD) to boost its impact on the generated left singular vectors. As a result, the similarities among documents in the same category are increased. Furthermore, an adaptive scaling strategy is designed to better utilize the hierarchical structure of categories. Experimental results show that the proposed approach is able to significantly improve the performance of hierarchical text categorization. Xiaojun Quan, Enhong Chen, Qiming Luo, Hui Xiong 0001 |
SIGIR | 1 |