VLDB 2026 Research / reviewers in the wild / expert
Xuejie Zhang 0002
dblp:68/3522-2
· DBLP profile ↗
115ranked-venue papers
0as first author
87since 2021 · last 2026
0000-0002-6591-0916ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 16 since 2021Systems, architecture and hardware · 11 · 6 since 2021Databases, data management, data science and information retrieval · 10 · 9 since 2021Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAPO: Self-Adaptive Process Optimization Makes Small Reasoners StrongerabstractExisting self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC) process supervision further exacerbates the difficulty in mitigating the gap. Motivated by the Error-Related Negativity (ERN), which the reasoner can localize error following incorrect decisions, guiding rapid adjustments, we propose a Self-Adaptive Process Optimization (SAPO) method for self-improvement in Small Language Models (SLMs). SAPO adaptively and efficiently introduces process supervision signals by actively minimizing the reasoner-verifier gap rather than relying on inefficient MC estimations. Extensive experiments demonstrate that the proposed method outperforms most existing self-evolution methods on two challenging task types: mathematics and code. Additionally, to further investigate SAPO's impact on verifier performance, this work introduces two new benchmarks for process reward models in both mathematical and coding tasks. Kaiyuan Chen 0002, Guangmin Zheng 0001, Jin Wang 0008, Xiaobing Zhou, Xuejie Zhang 0002 |
AAAI | 5 |
| 2026 | Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningabstractLarge reasoning models (LRMs) improve performance at test time by thinking longer, but this often leads to overthinking and high computational cost. To address this, recent reinforcement learning (RL) methods adopt outcome-level rewards, such as rule- or prompt-based signals, that favor shorter correct reasoning paths but often overlook reasoning quality. While such rewards neglect intermediate reasoning, dense supervision from process reward models (PRMs) has proven more effective in promoting coherent and high-quality reasoning. However, static PRM supervision introduces two challenges: reward hacking, since fixed rewards poorly capture global reasoning objectives, and the high training cost of obtaining dense reward labels at scale. To overcome these issues, we propose Step Group Relative Policy Optimization (Step-GRPO), a GRPO-based method that integrates step-level PRM signals into sparse trajectory-level feedback, avoiding costly step-level supervision while improving reasoning quality beyond accuracy. In addition, Step-GRPO employs a step-attention mechanism that captures inter-step dependencies and emphasizes critical reasoning steps, effectively mitigating reward hacking. We apply Step-GRPO to train large language models and observe consistent gains in reasoning quality, accuracy, and shorter reasoning traces across multiple math benchmarks, outperforming reinforcement learning baselines at substantially lower cost. Notably, the proposed model achieves 36.7 percent accuracy on AIME 2024 with 11,000 training samples and a training cost of 38 US dollars, surpassing baselines that require over 1,000 US dollars and more than 40,000 samples, demonstrating strong cost-effectiveness and scalability. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
AAAI | 4 |
| 2026 | LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models
Tiesunlong Shen, Rui Mao 0010, Jin Wang 0008, Heming Sun, Jian Zhang 0087, Xuejie Zhang 0002, Erik Cambria |
AAAI | 6 |
| 2026 | Syntax-aware question generation through dependency relations-guided attention
Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Self-verified user simulator via code-based interpretation in task-oriented dialogues
Xiang Luo 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Towards privacy-preserving and communication-efficient federated distillation
Xinge Ma, Jin Wang 0008, Xuejie Zhang 0002 |
Expert Syst. Appl. | 3 |
| 2026 | Fairness-efficiency tradeoffs in multiresource allocation for cloud-edge collaborative computing
Xiaobo Lin, Weidong Li 0002, Xuejie Zhang 0002 |
Future Gener. Comput. Syst. | 4 |
| 2026 | Cloud-edge collaborative task offloading and resource allocation based on mobile computility
Qian Su, Weidong Li 0002, Guangqin Hu, Xuejie Zhang 0002 |
Future Gener. Comput. Syst. | 5 |
| 2026 | EfficientLoRA: Rethinking the efficiency of low-rank adaptation in pre-trained language models
Shitong Cao, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Neural Networks | 3 |
| 2026 | Language-dominated fusion and self-distillation for multimodal sentiment analysis with incomplete modalities
Jin Wang 0008, Xuejie Zhang 0002 |
Pattern Recognit. | 3 |
| 2026 | Test-Time Domain-Agnostic Meta-Prompt Learning for Multi-Source Few-Shot Domain Adaptation
Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Sample-aware Adaptive Structured Pruning for Large Language ModelsabstractLarge language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deployment. Structured pruning can effectively reduce the resource demands for deployment by removing redundant model parameters. However, the randomly selected calibration data and fixed single importance estimation metrics in existing structured pruning methods lead to degraded performance of pruned models. This study introduces AdaPruner, a sample-aware adaptive structured pruning framework for LLMs, aiming to optimize the calibration data and importance estimation metrics in the structured pruning process. Specifically, AdaPruner effectively removes redundant parameters from LLMs by constructing a structured pruning solution space and then employing Bayesian optimization to adaptively search for the optimal calibration data and importance estimation metrics. Experimental results show that the AdaPruner outperforms existing structured pruning methods on a family of LLMs with varying pruning ratios, demonstrating its applicability and robustness. Remarkably, at a 20% pruning ratio, the model pruned with AdaPruner maintains 97% of the performance of the unpruned model. Jun Kong 0003, Xinge Ma, Jin Wang 0008, Xuejie Zhang 0002 |
AAAI | 4 |
| 2025 | Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain AdaptationabstractConventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptation (UMFDA) schema. It belongs to a decentralized edge collaborative learning in the edge-side models that must maintain a low computational load. And only a limited amount of annotations in source domain data is provided, with most of the data being unannotated. Further, this paper proposes a vision-aware multimodal prompt tuning framework (VAMP) under the decentralized schema, where the vision-aware prompt guides the text domain-specific prompt to maintain semantic discriminability and perceive the domain information. The cross-modal semantic and domain distribution alignment losses optimize each edge-side model, while text classifier consistency and semantic diversity losses promote collaborative learning among edge-side models. Extensive experiments were conducted on OfficeHome and DomainNet datasets to demonstrate the effectiveness of the proposed VAMP in the UMFDA, which outperformed the previous prompt tuning methods. Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 5 |
| 2025 | Data-Free Black-Box Federated Learning via Zeroth-Order Gradient EstimationabstractFederated learning (FL) enables decentralized clients to collaboratively train a global model under the orchestration of a central server without exposing their individual data. However, the iterative exchange of model parameters between the server and clients imposes heavy communication burdens, risks potential privacy leakage, and even precludes collaboration among heterogeneous clients. Distillation-based FL tackles these challenges by exchanging low-dimensional model outputs rather than model parameters, yet it highly relies on a task-relevant auxiliary dataset that is often not available in practice. Data-free FL attempts to overcome this limitation by training a server-side generator to directly synthesize task-specific data samples for knowledge transfer. However, the update rule of the generator requires clients to share on-device models for white-box access, which greatly compromises the advantages of distillation-based FL. This motivates us to explore a data-free and black-box FL framework via Zeroth-order Gradient Estimation (FedZGE), which estimates the gradients after flowing through on-device models in a black-box optimization manner to complete the training of the generator in terms of fidelity, transferability, diversity, and equilibrium, without involving any auxiliary data or sharing any model parameters, thus combining the advantages of both distillation-based FL and data-free FL. Experiments on large-scale image classification datasets and network architectures demonstrate the superiority of FedZGE in terms of data heterogeneity, model heterogeneity, communication efficiency, and privacy protection. Xinge Ma, Jin Wang 0008, Xuejie Zhang 0002 |
AAAI | 3 |
| 2025 | Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian PerspectiveabstractCurrent neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods affect pre-trained language models (PLMs) remains unclear. This study revisits the assumption that non-IID information enhances PLMs to achieve performance improvements from a Bayesian perspective, which unearths and integrates non-IID and IID features. Furthermore, we proposed a multi-attribute multi-grained framework for PLM adaptations (M2A), which combines multi-attribute and multi-grained views to mitigate uncertainty in a lightweight manner. We evaluate M2A through prevalent text-understanding datasets and demonstrate its superior performance, mainly when data are implicitly non-IID, and PLMs scale larger. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 5 |
| 2025 | Learning to Reason via Self-Iterative Process Feedback for Small Language ModelsabstractSmall language models (SLMs) are more efficient, cost-effective, and customizable than large language models (LLMs), though they often underperform in specific areas like reasoning. Past methods for enhancing SLMs’ reasoning, such as supervised fine-tuning and distillation, often depend on costly external signals, resulting in SLMs being overly confident with limited supervision signals, thus limiting their abilities. Therefore, this study enables SLMs to learn to reason from self-iterative feedback. By combining odds ratio preference optimization (ORPO), we fine-tune and align SLMs using positive and negative signals generated by themselves. Additionally, we introduce process supervision for rewards in preference alignment by sampling-based inference simulation and process reward models. Compared to Supervised Fine-Tuning (SFT), our method improves the performance of Gemma-2B by 12.43 (Acc) on GSM8K and 3.95 (Pass@1) on MBPP. Furthermore, the proposed method also demonstrated superior out-of-domain generalization capabilities on MMLU_Math and HumanEval. Kaiyuan Chen 0002, Jin Wang 0008, Xuejie Zhang 0002 |
COLING | 3 |
| 2025 | Topology-of-Question-Decomposition: Enhancing Large Language Models with Information Retrieval for Knowledge-Intensive TasksabstractLarge language models (LLMs) are increasingly deployed for general problem-solving across various domains yet remain constrained to chaining immediate reasoning steps and depending solely on parametric knowledge. Integrating an information retrieval system directly into the reasoning process of LLMs can improve answer accuracy but might disrupt the natural reasoning sequence. Consequently, LLMs may underperform in complex, knowledge-intensive tasks requiring multiple reasoning steps, extensive real-world knowledge, or critical initial decisions. To overcome these challenges, we introduce a novel framework, Topology-of-Question-Decomposition (ToQD), which activates retrieval only when necessary. Globally, ToQD guides LLMs in constructing a topology graph from the input question, each node representing a sub-question. Locally, ToQD employs self-verify inference to determine whether a sub-question should retrieve relevant documents, necessitate further decomposition, or directly provide an answer. Experiments demonstrate that ToQD achieves superior performance and robustness in complex, knowledge-intensive tasks, significantly enhancing system response efficiency. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 4 |
| 2025 | Reasoning with Trees: Faithful Question Answering over Knowledge GraphabstractRecent advancements in large language models (LLMs) have shown remarkable progress in reasoning capabilities, yet they still face challenges in complex, multi-step reasoning tasks. This study introduces Reasoning with Trees (RwT), a novel framework that synergistically integrates LLMs with knowledge graphs (KGs) to enhance reasoning performance and interpretability. RwT reformulates knowledge graph question answering (KGQA) as a discrete decision-making problem, leveraging Monte Carlo Tree Search (MCTS) to iteratively refine reasoning paths. This approach mirrors human-like reasoning by dynamically integrating the LLM’s internal knowledge with external KG information. We propose a real-data guided iteration technique to train an evaluation model that assesses action values, improving the efficiency of the MCTS process. Experimental results on two benchmark KGQA datasets demonstrate that RwT significantly outperforms existing state-of-the-art methods, with an average performance improvement of 9.81%. Notably, RwT achieves these improvements without requiring complete retraining of the LLM, offering a more efficient and adaptable approach to enhancing LLM reasoning capabilities. Tiesunlong Shen, Jin Wang 0008, Xuejie Zhang 0002, Erik Cambria |
COLING | 3 |
| 2025 | A Resource Allocation Method of Blockchain Network Based on Edge Computing
Qian Su, Longfei Bai, Guangqin Hu, Xuejie Zhang 0002 |
ICA3PP (3) | 5 |
| 2025 | Hop-level Direct Preference Optimization for Knowledge Graph Reasoning with TreesabstractRecent advancements in knowledge graph question answering (KGQA) have shown promise, yet existing methods often fail to align with human reasoning patterns. This study proposes HD-PORT (hop-level direct preference optimization for knowledge graph reasoning with trees), a novel approach that combines Monte Carlo Tree Search (MCTS) with hop-level direct preference optimization (HDPO) for KGQA tasks. MCTS simulates human-like reasoning by exploring multiple inference paths in knowledge graphs, generating interpretable reasoning chains and rich hop-level preferences. HDPO then leverages these preferences for model optimization, addressing the limitations of traditional supervised learning and existing DPO methods that focus on overall solution preferences. By optimizing preferences at each reasoning step, HD-PORT more closely mirrors human problem-solving strategies. Experimental results on benchmark datasets demonstrate that HD-PORT significantly outperforms state-of-the-art methods in both accuracy and interpretability, particularly for complex, multi-hop reasoning tasks. Tiesunlong Shen, Jin Wang 0008, Xuejie Zhang 0002, Erik Cambria |
ICASSP | 3 |
| 2025 | Dual-Path Contrastive Short Text Clustering with High-order Random WalkabstractIn recent years, several robust contrastive text clustering methods have been proposed. While these methods have achieved significant performances, two issues remain. First, the false negative problem is still not fully resolved, and the false positive issue also arises because all in-neighborhood and out-of-neighborhood samples are simply treated as positive and negative pairs, respectively. Second, these methods treat text representation learning and clustering as independent processes, leading to a performance gap. We propose a novel robust method called Dual-Path Contrastive Short Text Clustering (DCTC) to address these two issues. DCTC employs instance-level contrastive learning based on random walks at the representation learning level to progressively identify data pairs in a global rather than local manner, identifying in-neighborhood negatives and out-of-neighborhood positives. At the clustering level, DCTC performs cluster-level contrastive learning, jointly optimizing representation learning and cluster assignments, thereby enhancing clustering performance. DCTC achieves state-of-the-art results on 8 datasets, demonstrating its effectiveness. The code is available at https://github.com/2251821381/DCTC. Zhengzhong Zhu, Binjie Sun, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
ICASSP | 3 |
| 2025 | Enhanced Multimodal Chain-of-Thought with Visual Self-Contrastive DistillationabstractChain-of-thought (CoT) reasoning research has predominantly focused on language modality, neglecting the intricate interaction of multiple modalities crucial for real-world reasoning, such as visual question answering. Current methods primarily concentrate on modal conversion and feature fusion to enable language models to utilize CoT in a multimodal environment. However, these methods have inherent hallucination limitations. They tend to excessively rely on the prior knowledge of language unimodal and lack guidance in learning visual differences, often resulting in content incongruent with the images. This study introduces a novel approach, visual self-contrastive distillation (VSCD). The proposed method equalizes the status of language and vision modalities by freezing both encoders, enabling more balanced learning. Furthermore, we use contrastive decoding to enable the model to learn image differences during distillation, enhancing its understanding of visual nuances. The comprehensive experiment on the ScienceQA dataset demonstrates the superiority of the proposed VSCD method across various categories of multimodal CoT. Code and data are released at https://github.com/zgMin/VSCD. Guangmin Zheng 0001, Jun Kong 0003, Jin Wang 0008, Xuejie Zhang 0002 |
ICME | 4 |
| 2025 | Knowledge-Enhanced Question Generation Guided by Interrogative WordsabstractQuestion Generation (QG) focuses on creating relevant questions from a given context, but question-answering texts often contain substantial redundant or irrelevant information. The key to effective QG lies in identifying and selecting relevant phrases. For lengthy contexts, combining these discrete phrases into semantically coherent questions remains a significant challenge. To address the issue of information redundancy in long documents, this paper proposes a method that extracts key information from discrete documents to construct fine-grained text representations. This method incorporates inductive biases to support the Question Generation process. Additionally, the paper introduces an interrogative word type predictor that accurately identifies interrogative words and determines the domain of the answer, guiding the model in generating semantically aligned questions. This ensures that the generated questions are closely aligned with the answers and their respective context. The proposed method significantly reduces information redundancy and the reliance on annotated data. Experimental results on two benchmark data sets demonstrate that the proposed model achieves performance comparable to that of state-of-the-art QG methods. Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
IJCNN | 2 |
| 2025 | Dual Representation Space Optimization for Multi-Label Text ClassificationabstractMulti-label text classification (MLTC) presents significant challenges due to the need for accurate document representations and the effective modeling of complex label dependencies. Existing methods either underutilize label semantics for representation generation or face challenges in fully capturing inter-label dependencies using contrastive learning, often leading to suboptimal label predictions. To address these issues, we propose a novel end-to-end framework, Dual Representation Space Optimization (DRSO), for MLTC. DRSO tackles these challenges through two key components: a semantic-aware network that refines document representations by leveraging label semantics and an adaptive multi-label contrastive learning mechanism that captures inter-label dependencies to optimize label distributions in the prediction space. Extensive experiments on benchmark datasets demonstrate that DRSO outperforms state-of-the-art methods, showcasing its effectiveness in enhancing both representation quality and label prediction accuracy1. Binjie Sun, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
IJCNN | 2 |
| 2025 | Qwen-Gender: A Chain-of-Thought Based Multi-task Gender Bias Mitigation System
You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
NLPCC (4) | 5 |
| 2025 | EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection
Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (3) | 3 |
| 2025 | Flow-guided Direct Preference Optimization for Knowledge Graph Reasoning with TreesabstractRecent advancements in knowledge graph question answering (KGQA) have shown promise, yet existing methods often fail to align with human reasoning patterns that involve continuous reflection and refinement. This paper proposes FD-PORT (flow-guided direct preference optimization for knowledge graph reasoning with trees), a novel approach that combines Monte Carlo Tree Search (MCTS) with flow-guided direct preference optimization (FDPO) for KGQA tasks. MCTS simulates human-like reasoning by systematically exploring multiple inference paths in knowledge graphs, while FDPO transforms the search feedback into fine-grained training signals through flow balance conditions. Unlike traditional methods focusing on end-to-end training or sequence-level preferences, FD-PORT establishes flow consistency between any states along the reasoning chain, enabling robust multi-hop reasoning that adapts to local decisions and long-range dependencies. Experimental results on three benchmark datasets demonstrate that FD-PORT significantly outperforms state-of-the-art methods, achieving up to 50.6% improvements over GPT-4 on complex multi-hop reasoning tasks with a smaller open-source language model. The framework is advanced in maintaining diverse reasoning paths while ensuring answer quality, closely mirroring human problem-solving strategies. Tiesunlong Shen, Rui Mao 0010, Jin Wang 0008, Xuejie Zhang 0002, Erik Cambria |
SIGIR | 4 |
| 2025 | TopoDiff: Training-free image generation with topological layout control
Shitong Cao, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Expert Syst. Appl. | 2 |
| 2025 | Parameter-efficient online knowledge distillation for pretrained language models
Jin Wang 0008, Xuejie Zhang 0002 |
Expert Syst. Appl. | 3 |
| 2025 | Dual-dimensional contrastive learning for incomplete multi-view clustering
Zhengzhong Zhu, Chujun Pu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Neurocomputing | 3 |
| 2025 | Disentangled feature graph for Hierarchical Text Classification
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Inf. Process. Manag. | 2 |
| 2025 | Knowledge distillation via adaptive meta-learning for graph neural network
Tiesunlong Shen, Jin Wang 0008, Xuejie Zhang 0002 |
Inf. Sci. | 3 |
| 2025 | Heterogeneous federated distillation with mutual information maximization for medical relation extraction
Jin Wang 0008, Jiaxu Dao, You Zhang 0002, Dan Xu 0001, Xuejie Zhang 0002 |
Inf. Sci. | 5 |
| 2025 | Feature disentanglement, selection, and reaggregation method for multi-task learning
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Knowl. Inf. Syst. | 2 |
| 2025 | Distilling vision-language pre-training models with modality-specific meta-learning
Xinge Ma, Jin Wang 0008, Xuejie Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Syntax-Enhanced Pretrained Language Models for Aspect-Level Sentiment ClassificationabstractThe main challenge of aspect-level sentiment classification (ASC) is associating target aspect terms with relevant contextual words. Existing methods improve ASC performance by incorporating syntactic dependencies through a graph convolution layer on top of BERT. However, these approaches often assign a fixed weight to edges with the same dependency type, overlooking the contextual nuances these dependencies can convey. To address this, we propose syntax-enhanced BERT (SE-BERT), which integrates syntactic distance embeddings, a syntax-enhanced transformer, aspect-specific masking, and a sentiment classification layer. SE-BERT advances previous methods in two main ways. First, SE-BERT enables dynamic weighting for edges with the same dependency type based on the source node's part-of-speech (POS) tags, which enhances graph propagation accuracy and enables more precise edge associations. Second, rather than adding additional graph convolutional layers on top of BERT, SE-BERT replaces BERT's final Transformer layers with the proposed SE-Transformer. This approach directly encodes syntactic information into the attention distribution and word representation within scaled dot-product attention. The model can be initialized from a pretrained checkpoint and fine-tuned for downstream tasks without requiring additional training parameters. Experimental results on five benchmark datasets demonstrate that SE-BERT outperforms existing ASC methods. Jin Wang 0008, Lung-Hao Lee, Xuejie Zhang 0002 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Personalized LoRA for Human-Centered Text UnderstandingabstractEffectively and efficiently adapting a pre-trained language model (PLM) for human-centered text understanding (HCTU) is challenging since user tokens are million-level in most personalized applications and do not have concrete explicit semantics. A standard and parameter-efficient approach (e.g., LoRA) necessitates memorizing numerous suits of adapters for each user. In this work, we introduce a personalized LoRA (PLoRA) with a plug-and-play (PnP) framework for the HCTU task. PLoRA is effective, parameter-efficient, and dynamically deploying in PLMs. Moreover, a personalized dropout and a mutual information maximizing strategies are adopted and hence the proposed PLoRA can be well adapted to few/zero-shot learning scenarios for the cold-start issue. Experiments conducted on four benchmark datasets show that the proposed method outperforms existing methods in full/few/zero-shot learning scenarios for the HCTU task, even though it has fewer trainable parameters. For reproducibility, the code for this paper is available at: https://github.com/yoyo-yun/PLoRA. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 5 |
| 2024 | Zero-Shot Cross-Domain Dialogue State Tracking via Dual Low-Rank AdaptationabstractZero-shot dialogue state tracking (DST) seeks to enable dialogue systems to transition to unfamiliar domains without manual annotation or extensive retraining.Prior research has approached this objective by embedding prompts into language models (LMs).Common methodologies include integrating prompts at the input layer or introducing learnable variables at each transformer layer.Nonetheless, each strategy exhibits inherent limitations.Prompts integrated at the input layer risk underutilization, with their impact potentially diminishing across successive transformer layers.Conversely, the addition of learnable variables to each layer can complicate the training process and increase inference latency.To tackle the issues mentioned above, this paper proposes Dual Low-Rank Adaptation (DualLoRA), a plug-and-play architecture designed for zero-shot DST.DualLoRA incorporates two distinct Low-Rank Adaptation (LoRA) components, targeting both dialogue context processing and prompt optimization, to ensure the comprehensive influence of prompts throughout the transformer model layers.This is achieved without incurring additional inference latency, showcasing an efficient integration into existing architectures.Through rigorous evaluation on the MultiWOZ and SGD datasets, DualLoRA demonstrates notable improvements across multiple domains, outperforming traditional baseline methods in zeroshot settings.Our code is accessible at Xiang Luo 0003, Zhiwen Tang, Jin Wang 0008, Xuejie Zhang 0002 |
ACL (1) | 4 |
| 2024 | DuetSim: Building User Simulator with Dual Large Language Models for Task-Oriented DialoguesabstractUser Simulators play a pivotal role in training and evaluating task-oriented dialogue systems. Traditional user simulators typically rely on human-engineered agendas, resulting in generated responses that often lack diversity and spontaneity. Although large language models (LLMs) exhibit a remarkable capacity for generating coherent and contextually appropriate utterances, they may fall short when tasked with generating responses that effectively guide users towards their goals, particularly in dialogues with intricate constraints and requirements. This paper introduces DuetSim, a novel framework designed to address the intricate demands of task-oriented dialogues by leveraging LLMs. DuetSim stands apart from conventional approaches by employing two LLMs in tandem: one dedicated to response generation and the other focused on verification. This dual LLM approach empowers DuetSim to produce responses that not only exhibit diversity but also demonstrate accuracy and are preferred by human users. We validate the efficacy of our method through extensive experiments conducted on the MultiWOZ dataset, highlighting improvements in response quality and correctness, largely attributed to the incorporation of the second LLM. Xiang Luo 0003, Zhiwen Tang, Jin Wang 0008, Xuejie Zhang 0002 |
LREC/COLING | 4 |
| 2024 | SoftMCL: Soft Momentum Contrastive Learning for Fine-grained Sentiment-aware Pre-trainingabstractThe pre-training for language models captures general language understanding but fails to distinguish the affective impact of a particular context to a specific word. Recent works have sought to introduce contrastive learning (CL) for sentiment-aware pre-training in acquiring affective information. Nevertheless, these methods present two significant limitations. First, the compatibility of the GPU memory often limits the number of negative samples, hindering the opportunities to learn good representations. In addition, using only a few sentiment polarities as hard labels, e.g., positive, neutral, and negative, to supervise CL will force all representations to converge to a few points, leading to the issue of latent space collapse. This study proposes a soft momentum contrastive learning (SoftMCL) for fine-grained sentiment-aware pre-training. Instead of hard labels, we introduce valence ratings as soft-label supervision for CL to fine-grained measure the sentiment similarities between samples. The proposed SoftMCL conducts CL on both the word- and sentence-level to enhance the model’s ability to learn affective information. A momentum queue was introduced to expand the contrastive samples, allowing storing and involving more negatives to overcome the limitations of hardware platforms. Extensive experiments were conducted on four different sentiment-related tasks, which demonstrates the effectiveness of the proposed SoftMCL method. The code and data of the proposed SoftMCL is available at: https://www.github.com/wangjin0818/SoftMCL/. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
LREC/COLING | 3 |
| 2024 | Improving Personalized Sentiment Representation with Knowledge-enhanced and Parameter-efficient Layer NormalizationabstractExisting studies on personalized sentiment classification consider a document review as an overall text unit and incorporate backgrounds (i.e., user and product information) to learn sentiment representation. However, it is difficult when these methods meet the current pretrained language models (PLMs) owing to quadratic costs that increase with text length and heterogeneous mixes of randomly initialized background information and textual information initialized from well-pretrained checkpoints during information incorporation. To address these problems, we propose a knowledge-enhanced and parameter-efficient layer normalization (E2LN) for efficient and effective review modeling via leveraging LN in transformer structures. Initially, a knowledge base is introduced that stores well-pretrained checkpoints, structured text information, and background information. Based on such a knowledge base, the ability of LN can be magnified as being a crucial component of transformer structure and then improve the performance of PLMs in downstream tasks. Moreover, the proposed E2LN can make PLMs capable of modeling long document reviews and incorporating background information with parameter-efficient fine-tuning and knowledge injecting. Extensive experimental results were obtained for three document-level sentiment classification benchmark datasets. By comparing the results, the effectiveness and efficiency of the proposed model was demonstrated. Code and Data are released at https://github.com/yoyo-yun/E2LN. You Zhang 0002, Jin Wang 0008, Liang-Chih Yu, Dan Xu 0001, Xuejie Zhang 0002 |
LREC/COLING | 5 |
| 2024 | Enhancing Semantics in Multimodal Chain of Thought via Soft Negative SamplingabstractChain of thought (CoT) has proven useful for problems requiring complex reasoning. Many of these problems are both textual and multimodal. Given the inputs in different modalities, a model generates a rationale and then uses it to answer a question. Because of the hallucination issue, the generated soft negative rationales with high textual quality but illogical semantics do not always help improve answer accuracy. This study proposes a rationale generation method using soft negative sampling (SNSE-CoT) to mitigate hallucinations in multimodal CoT. Five methods were applied to generate soft negative samples that shared highly similar text but had different semantics from the original. Bidirectional margin loss (BML) was applied to introduce them into the traditional contrastive learning framework that involves only positive and negative samples. Extensive experiments on the ScienceQA dataset demonstrated the effectiveness of the proposed method. Code and data are released at https://github.com/zgMin/SNSE-CoT. Guangmin Zheng 0001, Jin Wang 0008, Xiaobing Zhou, Xuejie Zhang 0002 |
LREC/COLING | 4 |
| 2024 | Decoupling Control in Text-to-Image Diffusion Models
Shitong Cao, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
ICIC (7) | 2 |
| 2024 | Learning Defendant-aware Label Representation for Multi-Defendant Charge PredictionabstractAutomatic charge prediction based on deep learning methods is a crucial task in legal judgment prediction, aiming to predict the charges based on the fact description for a criminal case. While existing methods focus on multi-class cases with a single defendant, they fail to account for situations involving multiple defendants and labels, limiting their real-world application. To address these limitations, we propose a multi-defendant charge prediction approach that learns defendant-aware label representations (DLR). To handle complex circumstances for diverse defendants in a case, we extract defendant-specific representation by a machine reading comprehension approach via prompting the defendant’s name. In comparison with traditional text classifications that use discrete one-hot label representations, labels in charge predictions require clear definitions such as textual descriptions for determining the exact classified target. Therefore, we resort to a label description encoder to facilitate the charge predictions via understanding defendant-specific representations. Accordingly, we empower an efficient low-rank adaption module as a feature fuser that incorporates dependent-specific representations into label encoders. The proposed method is evaluated on both multi- and single-dependent-based charge prediction datasets, showing its comparable performances in real-world scenarios. The codes and collected datasets for our study are available at: https://github.com/cy330874054/LDLRMDCP. You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
IJCNN | 5 |
| 2024 | Hierarchical Differential Amplifier Contrastive Learning for Semi-supervised Extractive SummarizationabstractExtractive summarization aims to generate summaries by extracting and concatenating critical information from long documents. However, existing extractive summarization methods typically rely on large-scale labeled datasets, which are expensive and time-consuming. In addition, due to the inherent class imbalance problem in extractive summarization, traditional approaches such as resampling make it difficult to address this challenge effectively. To address these issues, we treat extractive summarization as a class imbalance problem and propose a method called HDCSUM. Specifically, we first employ consistency-training and pseudo-labeling strategies to fully use limited labeled data and many unlabeled data. Additionally, HD-CSUM can amplify the idiosyncratic information of each sentence and focus more on the semantic differences by jointly utilizing differential amplifiers and contrastive learning. We introduce a weighted cross-entropy loss function to further address the class imbalance problem. Extensive experiments on the challenging and widely used CNN/Daily Mail and BBC XSum benchmark datasets show that our method outperforms state-of-the-art semi-supervised methods. Our code is available at GitHub1. Jiankuo Li, Xuejie Zhang 0002, Jin Wang 0008, Shitong Cao, Xiaobing Zhou |
IJCNN | 2 |
| 2024 | LoRA-Enhanced Language Alignments for Robust Code-Mixed Text RepresentationabstractThe utilization of code-mixed texts allows individuals the opportunity to express sentiments flexibly in international and multilingual contexts. However, the diversity of languages and pragmatic writing styles can lead to semantic shifts at both word and sentence levels, resulting in a degraded comprehension of code-mixed texts by machines. To tackle this issue, we propose a method for language alignments, which aligns both word- and sentence-level semantic representations via a low-rank injection (LoRI) and a data augmentation strategy (DA), dubbed LoRIDA. LoRI integrates linguistic features into textual representations as a feature fusion mechanism. To further bridge the gaps between sentence-level semantics, we augment code-mixed data into individual source languages and apply a knowledge distillation method for joint alignments. We evaluate the performance of the proposed method on four code-mixed sentiment analysis datasets, demonstrating its superiority over existing methods. Our code is publicly available at https://github.com/linsongisgood/LELA. Xuqiao Ran, You Zhang 0002, Jin Wang 0008, Dan Xu 0001, Xuejie Zhang 0002 |
IJCNN | 5 |
| 2024 | Mathematical Reasoning via Multi-step Self Questioning and Answering for Small Language Models
Kaiyuan Chen 0002, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (4) | 3 |
| 2024 | PROMPTIST: Automated Prompt Optimization for Text-to-Image Synthesis
Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (2) | 3 |
| 2024 | Chinese Metaphor Recognition Using a Multi-stage Prompting Large Language Model
Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (5) | 3 |
| 2024 | Disagreement Evaluation of Solutions for Math Word Problem
Yehui Xu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
ECML/PKDD (5) | 2 |
| 2024 | Fair multiresource allocation with access constraint in cloud-edge systems
Guangqin Hu, Weidong Li 0002, Xuejie Zhang 0002 |
Future Gener. Comput. Syst. | 4 |
| 2024 | Debiased momentum contrastive learning for multimodal video similarity measures
Kuanghong Liu, Jin Wang 0008, Xuejie Zhang 0002 |
Neurocomputing | 3 |
| 2024 | A machine reading comprehension model with counterfactual contrastive learning for emotion-cause pair extraction
Hanjie Mai, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Knowl. Inf. Syst. | 2 |
| 2024 | Hybrid-Mode tracker with online SA-LSTM updater
Hongsheng Zheng, Yaqing Hu, Xuejie Zhang 0002 |
Neural Comput. Appl. | 4 |
| 2024 | Joint contrastive learning for prompt-based few-shot language learners
Zhengzhong Zhu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Neural Comput. Appl. | 2 |
| 2024 | Layerwised multimodal knowledge distillation for vision-language pretrained model
Jin Wang 0008, Dawei Liao, You Zhang 0002, Dan Xu 0001, Xuejie Zhang 0002 |
Neural Networks | 5 |
| 2024 | Encoding Syntactic Information into Transformers for Aspect-Based Sentiment Triplet ExtractionabstractAspect-based sentiment triplet extraction(ASTE) aims to extract triplets consisting of aspect terms and their associated opinion terms and sentiment polarities from sentences, a relatively new and challenging subtask of aspect-based sentiment analysis (ABSA). Previous studies have used either pipeline models or unified tagging schema models. These models ignore the syntactic relationships between the aspect and its corresponding opinion words, which leads them to mistakenly focus on syntactically unrelated words. One feasible option is to use a graph convolution network (GCN) to exploit syntactic information by propagating the representation from the opinion words to the aspect. However, such a method considers all syntactic dependencies to be of the same type and thus may still incorrectly associate unrelated words to the target aspect through the iterations of graph convolutional propagation. Herein, a syntax-aware transformer (SA-Transformer) is proposed to extend the GCN strategy by fully exploiting the dependency types of edges to block inappropriate propagation. The proposed approach can obtain different representations and weights even for edges with the same dependency type according to their adjacent dependency type of edges. Instead of using a GCN layer, we used anL-layer SA transformer to encode syntactic information in the word-pair representation to improve performance. Experimental results on four benchmark datasets show that the proposed model outperforms various previous models for ASTE. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Adaptive Ensemble Self-Distillation With Consistent Gradients for Fast Inference of Pretrained Language ModelsabstractConditional computation algorithms, e.g., the early exiting (EE) strategy, can accelerate the inference of pretrained language models (PLMs) by exiting shallow layers without calculating the entire model. In addition to the adaptive inference of EE prediction for downstream tasks, self-distillation (SD) can encourage EE classifiers to mimic the behavior of the final classifier to enhance their representation capacity. However, the gradients from different tasks of EE classifiers will conflict and cancel one another out. The parameters of the backbone and some EE classifiers will be implicitly prevented from updating. Moreover, if the semantic gap between the final classifier and EE classifiers is significant, the EE classifier's performance will decrease. That is, the final classifier would not be the best choice to enhance the performance of the EE classifier. This study proposed an early exiting strategy with adaptive ensemble self-distillation and consistent gradients for the inference acceleration of PLMs. To mitigate gradient conflicts, we orthogonally projected the distillation loss's backpropagated gradient onto the classification loss's normal plane. Instead of directly using the final classifier as a single-teacher for self-distillation, we dynamically assemble different adaptive teachers for different EE classifiers according to the learning abilities of the EE classifiers. The accumulative decision was drawn for adaptive inference to make accurate and reliable predictions. Experimental results show that the proposed model outperforms existing models with the same speed-up ratio and effectively balances model performance and inference time. Jun Kong 0003, Jin Wang 0008, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Decoupled Knowledge Embedded Graph Convolutional Network for Skeleton-Based Human Action RecognitionabstractSkeleton-based action recognition has broad prospects owing to the fact that skeleton data is more robust to scene noise and camera view changes. Recently, researchers mainly aim to explore deep-learning feature engineering with competitive recognition accuracy for skeleton actions. However, a high-performance recognition network is usually stacked by complex feature extraction modules introducing massive computational costs. In this work, we designed a powerful and universal action knowledge distillation paradigm based on decoupled knowledge distillation for transferring action knowledge from heavy teachers to lightweight students more robustly. We constructed a network architecture space consisting of the shrinking versions of outdated 2s-AGCN and searched for several robust students. On this basis, this paradigm is further developed into a powerful decoupled knowledge embedded graph convolutional network (DKE-GCN), which outperforms the teacher significantly on three public datasets and achieves the state-of-the-art. In addition, a light-DKE-GCN is designed to achieve comparable performance with teacher with 16× less parameters, 26× less FLOPs and 8× FPS. Hao Zhang 0110, Xuejie Zhang 0002, Dan Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Primal-Dual-Based Computation Offloading Method for Energy-Aware Cloud-Edge CollaborationabstractIn the context of the Internet of Things (IoT), resource-constrained mobile edge computing (MEC) can no longer fully meet the needs of the rapidly growing number of mobile users; hence, cloud-edge collaborative computing has been developed. This paper focuses on the total energy consumption of the system and heterogeneity of scenarios, and a collaborative cloud-edge computation offloading approach with near real-time decision making is proposed. First, a general cloud-edge collaborative computation offloading model is abstracted from typical applications, and the energy consumption for edge and cloud offloading is calculated separately by considering both transmission and computational energy consumption. The problem is formulated as an integer linear program (ILP) with multidimensional resource constraints and is proven to be NP-hard. Then, a novel primal-dual computation offloading (PDCO) algorithm is designed to make near real-time offloading decisions one by one based on the sequential arrival of task requests. The approximation ratio of PDCO is derived through the weak duality property and the price-resource increment relationship. The experimental results show that under the guidance of total cost influenced by marginal prices, PDCO not only avoids blindly making offloading decisions but also effectively alleviates the shortage of resources on edge servers (ESs), approaching the optimal performance in terms of total energy consumption and resource utilization. Qian Su, Weidong Li 0002, Xuejie Zhang 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Multimodality Self-distillation for Fast Inference of Vision and Language Pretrained ModelsabstractThe computational cost of the vision and language pretrained models (VL-PTMs) limits their deployment in resource-constrained devices that require low latency. One existing solution is to apply the early exiting (EE) strategy to accelerate the inference. This technique can force model prediction using only a few former transformer layers. However, these former layers behave differently with the final classifier, inevitably resulting in performance decline. To counter such limitation, self-distillation has been commonly introduced to enhance the representation abilities of the EE classifiers. This results in a semantic gap since EE classifiers are directly trained to mimic the outputs of the final classifier without access to the modality-specific behaviors. This study proposes a multimodality self-distillation method for the fast inference of VL-PTMs. To fill the semantic gap between modalities, we split the multimodalities into separate modalities and added them as extra inputs to encourage the effective distillation of each modality. Furthermore, the mean squared error (MSE) is introduced to minimize the distance of feature maps and further enhance the representation ability of the EE classifiers. Experiments show that the proposed method outperforms the previous EE strategies with the same inference time, and performs competitively even if the model exited very early. Jun Kong 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Learning to Memorize Entailment and Discourse Relations for Persona-Consistent DialoguesabstractMaintaining engagement and consistency is particularly important in dialogue systems. Existing works have improved the performance of dialogue systems by intentionally learning interlocutor personas with sophisticated network structures. One issue with this approach is that it requires more personal corpora with annotations. Additionally, these models typically perform the next utterance prediction to generate a response but neglect the discourse coherence in the entire conversation. To address these issues, this study proposes a method of learning to memorize entailment and discourse relations for persona-consistent dialogue tasks. Entailment text pairs in natural language inference dataset were applied to learn latent entailment relations as external memories by premise-to-hypothesis generation task. Furthermore, an internal memory with a similar architecture was applied to the discourse information in the dialogue. Placing orthogonality restrictions on these two memory spaces ensures that the latent entailment relations remain dialogue-independent. Both memories collaborate to obtain entailment and discourse representation for the generation, allowing a deeper understanding of both consistency and coherence. Experiments on two large public datasets, PersonaChat and DSTC7-AVSD, demonstrated the effectiveness of the proposed method. Both automatic and human evaluations indicate that the proposed model outperforms several strong baselines in terms of both persona consistency and response coherence. Our source code is availabled at https://github.com/Chenrj233/LMEDR. Ruijun Chen 0001, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
AAAI | 4 |
| 2023 | FedID: Federated Interactive Distillation for Large-Scale Pretraining Language ModelsabstractThe growing concerns and regulations surrounding the protection of user data privacy have necessitated decentralized training paradigms.To this end, federated learning (FL) is widely studied in user-related natural language processing (NLP).However, it suffers from several critical limitations including extensive communication overhead, inability to handle heterogeneity, and vulnerability to white-box inference attacks.Federated distillation (FD) is proposed to alleviate these limitations, but its performance is faded by confirmation bias.To tackle this issue, we propose Federated Interactive Distillation (FedID), which utilizes a small amount of labeled data retained by the server to further rectify the local models during knowledge transfer.Additionally, based on the GLUE benchmark, we develop a benchmarking framework across multiple tasks with diverse data distributions to contribute to the research of FD in NLP community.Experiments show that our proposed Fe-dID framework achieves the best results in homogeneous and heterogeneous federated scenarios.The code for this paper is available at: https://github.com/maxinge8698/FedID. Xinge Ma, Jiangming Liu, Jin Wang 0008, Xuejie Zhang 0002 |
EMNLP | 4 |
| 2023 | Matrix Contrastive Learning for Short Text Clustering
Zhengzhong Zhu, Jiankuo Li, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
ICONIP (7) | 3 |
| 2023 | Semantic Candidate Retrieval for Few-Shot Entity Linking
Jianyong Chen, Jiangming Liu, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (3) | 4 |
| 2023 | Entity-Related Unsupervised Pretraining with Visual Prompts for Multimodal Aspect-Based Sentiment Analysis
Kuanghong Liu, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (2) | 3 |
| 2023 | An adversarial training-based mutual information constraint method
Renyuan Liu, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Appl. Intell. | 2 |
| 2023 | Interactive capsule network for implicit sentiment analysis
Yanjun Qian, Jin Wang 0008, Xuejie Zhang 0002 |
Appl. Intell. | 4 |
| 2023 | Graphs get personal: learning representation with contextual pretraining for collaborative filtering
Tiesunlong Shen, You Zhang 0002, Jin Wang 0008, Xuejie Zhang 0002 |
Appl. Intell. | 4 |
| 2023 | Causal representation for few-shot text classification
Maoqin Yang, Xuejie Zhang 0002, Jin Wang 0008, Xiaobing Zhou |
Appl. Intell. | 2 |
| 2023 | Decoupled variational autoencoder with interactive attention for affective text generation
Ruijun Chen 0001, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Multiresource fair allocation with time window constraints
Weidong Li 0002, Xuejie Zhang 0002 |
J. Supercomput. | 3 |
| 2023 | Fidelity-driven Optimization Reconstruction and Details Preserving Guided Fusion for Multi-Modality Medical ImageabstractBy integrating effective features of multi-modality medical images to provide richer information, multi-modality medical image fusion has been substantially used in computer-aided diagnosis applications. However, many existing fusion schemes do not consider how to eliminate the effects of the noise in source medical images and cannot provide enough details and textures for disease diagnosis. To address the problems above, we propose a new fidelity-driven optimization (FDO) reconstruction and details preserving guided-based fusion method for multi-modality medical images. To overcome the influence of noise in multi-modality medical images, a rank coefficient optimization method of low-rank approximation based on weighted mean curvature is proposed to reconstruct multi-modality medical image. Moreover, we propose an iterative detail preserving guided fusion (DPGF) method to integrate more textures and detail information of source multi-modality medical images, while ensuring high signal-to noise ratios. The experimental results show that the proposed method outperforms some of the state-of-the-art fusion methods. Specifically, the extensive experiments prove that our method has high robustness for noisy medical images, which also indicates the application prospects in diagnosis applications. Kangjian He, Xuejie Zhang 0002, Dan Xu 0001, Lisiqi Xie |
IEEE Trans. Multim. | 2 |
| 2022 | Accelerating Inference for Pretrained Language Models by Unified Multi-Perspective Early ExitingabstractConditional computation algorithms, such as the early exiting (EE) algorithm, can be applied to accelerate the inference of pretrained language models (PLMs) while maintaining competitive performance on resource-constrained devices. However, this approach is only applied to the vertical architecture to decide which layers should be used for inference. Conversely, the operation of the horizontal perspective is ignored, and the determination of which tokens in each layer should participate in the computation fails, leading to a high redundancy for adaptive inference. To address this limitation, a unified horizontal and vertical multi-perspective early exiting (MPEE) framework is proposed in this study to accelerate the inference of transformer-based models. Specifically, the vertical architecture uses recycling EE classifier memory and weighted self-distillation to enhance the performance of the EE classifiers. Then, the horizontal perspective uses recycling class attention memory to emphasize the informative tokens. Conversely, the tokens with less information are truncated by weighted fusion and isolated from the following computation. Based on this, both horizontal and vertical EE are unified to obtain a better tradeoff between performance and efficiency. Extensive experimental results show that MPEE can achieve higher acceleration inference with competent performance than existing competitive methods. Jun Kong 0003, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 4 |
| 2022 | Knowledge Distillation with Reptile Meta-Learning for Pretrained Language Model CompressionabstractThe billions, and sometimes even trillions, of parameters involved in pre-trained language models significantly hamper their deployment in resource-constrained devices and real-time applications. Knowledge distillation (KD) can transfer knowledge from the original model (i.e., teacher) into a compact model (i.e., student) to achieve model compression. However, previous KD methods have usually frozen the teacher and applied its immutable output feature maps as soft labels to guide the student’s training. Moreover, the goal of the teacher is to achieve the best performance on downstream tasks rather than knowledge transfer. Such a fixed architecture may limit the teacher’s teaching and student’s learning abilities. Herein, a knowledge distillation method with reptile meta-learning is proposed to facilitate the transfer of knowledge from the teacher to the student. The teacher can continuously meta-learn the student’s learning objective to adjust its parameters for maximizing the student’s performance throughout the distillation process. In this way, the teacher learns to teach, produces more suitable soft labels, and transfers more appropriate knowledge to the student, resulting in improved performance. Unlike previous KD using meta-learning, the proposed method only needs to calculate the first-order derivatives to update the teacher, leading to lower computational cost but better convergence. Extensive experiments on the GLUE benchmark show the competitive performance achieved by the proposed method. For reproducibility, the code for this paper is available at: https://github.com/maxinge8698/ReptileDistil. Xinge Ma, Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
COLING | 4 |
| 2022 | An Enhanced Key-utterance Interactive Model with Decouped Auxiliary Tasks for Multi-party Dialogue Reading ComprehensionabstractMulti-party dialogue machine reading comprehension (MRC) is more challenging than plain text MRC because it involves multiple speakers, more complex information flow interaction, and discourse structure. Previously most researchers focus on decoupling the speaker-aware and utterance-aware information to overcome such difficulties. Based on this, the self- and pseudo-self-supervised prediction auxiliary tasks on speakers and key-utterance are proposed. However, the information interaction among key-utterance, question, and dialogue context was ignored in these works, and there should also be a constraint between the two additional tasks. Herein, we proposed an enhanced key-utterance interaction model. It takes the key-utterance predicted by auxiliary task as prior information. Moreover, the co-attention mechanism is used to capture the critical information interaction among dialogue contexts, question, and key-utterance from the two perspectives of question-to-dialogue and dialogue-to-question, respectively. In addition, we introduced minimizing mutual information (MI) between the two auxiliary tasks to prevent mutual interference and overlap of information. Experimental results show that the proposed model achieves significant improvements than the dialogue MRC baseline models in Molweni and FriendsQA datasets. Xingyu Zhu 0006, Jin Wang 0008, Xuejie Zhang 0002 |
IJCNN | 3 |
| 2022 | An On-Device Machine Reading Comprehension Model with Adaptive Fast Inference
Fulai Nan, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (1) | 3 |
| 2022 | Hierarchical template transformer for fine-grained sentiment controllable generation
Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
Inf. Process. Manag. | 4 |
| 2022 | Hierarchical BERT with an adaptive fine-tuning strategy for document classification
Jun Kong 0003, Jin Wang 0008, Xuejie Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Contextual sentiment embeddings via bi-directional GRU language modelabstractCompared with conventional word embeddings, sentiment embeddings can distinguish words with similar contexts but opposite sentiment. They can be used to incorporate sentiment information from labeled corpora or lexicons by either end-to-end training or sentiment refinement. However, these methods present two major limitations. First, traditional approaches provide a fixed representation to each word but ignore the alternation of word meaning in different contexts. As a result, the polarity of a certain emotional word may vary with context, but will be assigned with a same representation. Another problem is the handling of out-of-vocabulary (OOV) or informal-writing sentiment words that would be assigned generic vectors (e.g., ). In addition, if affective words are not included in affective corpora or lexicons, they would be treated as neutral. Using such low-quality embeddings for building a neural model will reduce performance. This study proposes a training model of contextual sentiment embeddings. A stacked two-layer GRU model was used as the language model, simultaneously trained to incorporate semantic and sentiment information from labeled corpora and lexicons. To deal with OOV or informal-writing sentiment words, the WordPiece tokenizer was used to divide the text into subwords. The resulting model can be transferred to downstream applications by either feature extractor or fine-tuning. The results show that the proposed model can handle unseen or informal writing sentiment words and thus outperforms previously proposed methods. Jin Wang 0008, You Zhang 0002, Liang-Chih Yu, Xuejie Zhang 0002 |
Knowl. Based Syst. | 4 |
| 2022 | Explainable detection of adverse drug reaction with imbalanced data distributionabstractAnalysis of health-related texts can be used to detect adverse drug reactions (ADR). The greatest challenge for ADR detection lies in imbalanced data distributions where words related to ADR symptoms are often minority classes. As a result, trained models tend to converge to a point that strongly biases towards the majority class and then ignores the minority class. Since the most used cross-entropy criteria is an approximation to accuracy, the model focuses more readily on the majority class to achieve high accuracy. To address this issue, existing methods apply either oversampling or down-sampling strategies to balance the data distribution and exploit the most difficult samples of the minority class. However, increasing or reducing the number of individual tokens alone in sequence labeling tasks will result in the loss of the syntactic relations of the sentence. This paper proposes a weighted variant of conditional random field (CRF) for data-imbalanced sequence labeling tasks. Such a weighting strategy can alleviate data distribution imbalances between majority and minority classes. Instead of using softmax in the output layer, the CRF can capture the relationship of labels between tokens. The locally interpretable model-agnostic explanations (LIME) algorithm was applied to investigate performance differences between models with and without the weighted loss function. Experimental results on two different ADR tasks show that the proposed model outperforms previously proposed sequence labeling methods. Jin Wang 0008, Liang-Chih Yu, Xuejie Zhang 0002 |
PLoS Comput. Biol. | 3 |
| 2021 | Variational Autoencoder with Interactive Attention for Affective Text Generation
Ruijun Chen 0001, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (2) | 3 |
| 2021 | Accelerating Pretrained Language Model Inference Using Weighted Ensemble Self-distillation
Jun Kong 0003, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (1) | 3 |
| 2021 | Strategy-Proof Mechanism for Online Time-Varying Resource Allocation with Restart
Jixian Zhang 0003, Xuejie Zhang 0002, Weidong Li 0002 |
J. Grid Comput. | 3 |
| 2021 | Conciseness is better: Recurrent attention LSTM model for document-level sentiment analysis
You Zhang 0002, Jin Wang 0008, Xuejie Zhang 0002 |
Neurocomputing | 3 |
| 2021 | Learning sentiment sentence representation with multiview attention model
You Zhang 0002, Jin Wang 0008, Xuejie Zhang 0002 |
Inf. Sci. | 3 |
| 2021 | Personalized sentiment classification of customer reviews via an interactive attributes attention model
You Zhang 0002, Jin Wang 0008, Xuejie Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2020 | An online auction mechanism for time-varying multidimensional resource allocation in clouds
Jixian Zhang 0003, Xutao Yang, Xuejie Zhang 0002, Athanasios V. Vasilakos, Weidong Li 0002 |
Future Gener. Comput. Syst. | 4 |
| 2020 | Adversarial learning of sentiment word representations for sentiment analysis
Jin Wang 0008, Xuejie Zhang 0002 |
Inf. Sci. | 3 |
| 2020 | Pipelined Neural Networks for Phrase-Level Sentiment Intensity PredictionabstractLinguistic modifiers such as negators (e.g., not), intensifiers (e.g., very) and modals (e.g., would) are commonly used in expressing opinions. These modifiers play an important role in recognizing the sentiment intensity of multi-word phrases because they may lead to an intensity shift and polarity reversal for the words they modify. Appropriately modeling the effect of such modifiers on the intensity shift can greatly improve the performance of phrase-level sentiment intensity prediction. To this end, this paper proposes two neural network (NN) models organized in a pipelined fashion to determine 1) the intensity of individual words and 2) the shift weights of modifiers representing the degrees of intensity change for the words they modify. The intensity of a phrase can then be determined by combining the intensity of the constituent word and the shift weight of the modifier within the phrase. When measuring the word intensity, the first NN model introduces a hidden layer as a filter to select appropriate similar seed words in the prediction process. Automatic word intensity prediction can address the unknown intensities of words not covered in sentiment lexicons. In learning the modifier weights, the second NN model considers both the weights of individual modifiers and groups of modifiers to capture various intensity shift effects caused by them. Experiments on a SemEval-2016 dataset showed that the proposed method yielded better prediction performance for both single words and multi-word phrases. Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | Tree-Structured Regional CNN-LSTM Model for Dimensional Sentiment AnalysisabstractDimensional sentiment analysis aims to recognize continuous numerical values in multiple dimensions such as the valence-arousal (VA) space. Compared to the categorical approach that focuses on sentiment classification such as binary classification (i.e., positive and negative), the dimensional approach can provide a more fine-grained sentiment analysis. This article proposes a tree-structured regional CNN-LSTM model consisting of two parts: regional CNN and LSTM to predict the VA ratings of texts. Unlike a conventional CNN which considers a whole text as input, the proposed regional CNN uses a part of the text as a region, dividing an input text into several regions such that the useful affective information in each region can be extracted and weighted according to their contribution to the VA prediction. Such regional information is sequentially integrated across regions using LSTM for VA prediction. By combining the regional CNN and LSTM, both local (regional) information within sentences and long-distance dependencies across sentences can be considered in the prediction process. To further improve performance, a region division strategy is proposed to discover task-relevant phrases and clauses to incorporate structured information into VA prediction. Experimental results on different corpora show that the proposed method outperforms lexicon-, regression-, conventional NN and other structured NN methods proposed in previous studies. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Investigating Dynamic Routing in Tree-Structured LSTM for Sentiment AnalysisabstractJin Wang, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Visual Object Tracking via an Improved Lightweight Siamese Network
Xuejie Zhang 0002 |
PRCV (1) | 5 |
| 2019 | Approximation algorithm for the energy-aware profit maximizing problem in heterogeneous computing systems
Weidong Li 0002, Xi Liu 0002, Xiaobo Cai, Xuejie Zhang 0002 |
J. Parallel Distributed Comput. | 4 |
| 2019 | Multi-focus image fusion combining focus-region-level partition and pulse-coupled neural network
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie, Xin Jin 0005 |
Soft Comput. | 3 |
| 2018 | Multi-choice Virtual Machine Allocation with Time Windows in Cloud Computing
Jixian Zhang 0003, Xuejie Zhang 0002, Weidong Li 0002 |
GPC | 3 |
| 2018 | An online auction mechanism for cloud computing resource allocation and pricing based on user evaluation and cost
Jixian Zhang 0003, Xuejie Zhang 0002, Weidong Li 0002 |
Future Gener. Comput. Syst. | 3 |
| 2018 | Multi-focus: Focused region finding and multi-scale transform for image fusion
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie |
Neurocomputing | 3 |
| 2018 | Using a stacked residual LSTM model for sentiment intensity prediction
Jin Wang 0008, Xuejie Zhang 0002 |
Neurocomputing | 3 |
| 2018 | Refining Word Embeddings Using Intensity Scores for Sentiment AnalysisabstractWord embeddings that provide continuous low-dimensional vector representations of words have been extensively used for various natural language processing tasks. However, existing context-based word embeddings such as Word2vec and GloVe typically fail to capture sufficient sentiment information, which may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance. To tackle this problem, recent studies have suggested learning sentiment embeddings to incorporate the sentiment polarity (positive and negative) information from labeled corpora. This study adopts another strategy to learn sentiment embeddings. Instead of creating a new word embedding from labeled corpora, we propose a word vector refinement model to refine existing pretrained word vectors using real-valued sentiment intensity scores provided by sentiment lexicons. The idea of the refinement model is to improve each word vector such that it can be closer in the lexicon to both semantically and sentimentally similar words (i.e., those with similar intensity scores) and further away from sentimentally dissimilar words (i.e., those with dissimilar intensity scores). An obvious advantage of the proposed method is that it can be applied to any pretrained word embeddings. In addition, the intensity scores can provide more fine-grained (real-valued) sentiment information than binary polarity labels to guide the refinement process. Experimental results show that the proposed refinement model can improve both conventional word embeddings and previously proposed sentiment embeddings for binary, ternary, and fine-grained sentiment classification on the SemEval and Stanford Sentiment Treebank datasets. Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Strategy-Proof Mechanism for Provisioning and Allocation Virtual Machines in Heterogeneous CloudsabstractIn this paper, we address the problem of heterogeneous physical machines resource management (HPMRM); that is, providing and allocating multiple virtual machine (VM) instances from heterogeneous physical machines to maximize social welfare. Although existing allocation mechanisms allocate VMs to users through the single-mapping mechanism, such allocations cannot guarantee maximum social welfare or efficient utilization of multiple types of resources for cloud providers. Thus, we consider the multi-mapping mechanism, which permits mapping VMs allocated to one user to physical machines for VM provisioning and allocation. This can result in improved social welfare and lead to less resource fragmentation. We formulate the HPMRM problem in an auction-based setting, and design optimal and approximate mechanisms to solve it. In addition, we show that our proposed mechanism is strategy-proof; that is, our proposed mechanism drives the system into an equilibrium where no users have incentives to maximize their own profit by untruthfully reporting their requests. Furthermore, we analyze the approximation ratio of our proposed approximation algorithm. We also perform experiments to investigate the performance of our proposed approximation mechanism compared to the optimal mechanism. Experimental results demonstrate that our proposed approximation mechanism can obtain near optimal solutions and significantly improve allocation efficiency, while generating greater social welfare. Xi Liu 0002, Weidong Li 0002, Xuejie Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Refining Word Embeddings for Sentiment AnalysisabstractWord embeddings that can capture semantic and syntactic information from contexts have been extensively used for various natural language processing tasks.However, existing methods for learning contextbased word embeddings typically fail to capture sufficient sentiment information.This may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance.Therefore, this study proposes a word vector refinement model that can be applied to any pre-trained word vectors (e.g., Word2vec and GloVe).The refinement model is based on adjusting the vector representations of words such that they can be closer to both semantically and sentimentally similar words and further away from sentimentally dissimilar words.Experimental results show that the proposed method can improve conventional word embeddings and outperform previously proposed sentiment embeddings for both binary and fine-grained classification on Stanford Sentiment Treebank (SST). Liang-Chih Yu, Jin Wang 0008, K. Robert Lai, Xuejie Zhang 0002 |
EMNLP | 4 |
| 2017 | A Profit-Maximum Resource Allocation Approach for Mapreduce in Data Centers
Weidong Li 0002, Xi Liu 0002, Xuejie Zhang 0002 |
GPC | 4 |
| 2016 | Discrete Interior Search Algorithm for Multi-resource Fair Allocation in Heterogeneous Cloud Computing Systems
Xi Liu 0002, Weidong Li 0002, Xuejie Zhang 0002 |
ICIC (1) | 4 |
| 2016 | Building Chinese Affective Resources in Valence-Arousal Dimensions
Liang-Chih Yu, Lung-Hao Lee, Jin Wang 0008, Yunchao He, K. Robert Lai, Xuejie Zhang 0002 |
HLT-NAACL | 8 |
| 2016 | Locally weighted linear regression for cross-lingual valence-arousal prediction of affective words
Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
Neurocomputing | 4 |
| 2016 | Community-Based Weighted Graph Model for Valence-Arousal Prediction of Affective WordsabstractCompared to the categorical approach that represents affective states as several discrete classes (e.g., positive and negative), the dimensional approach represents affective states as continuous numerical values in multiple dimensions, such as the valence-arousal (VA) space, thus allowing for more fine-grained sentiment analysis. In building dimensional sentiment applications, affective lexicons with VA ratings are useful resources but are still very rare. Several semi-supervised methods such as the kernel method, linear regression, and the pagerank algorithm have been investigated to automatically determine the VA ratings of affective words from a set of semantically similar seed words. These methods suffer from two major limitations. First, they apply an equal weight to all seeds similar to an unseen word in predicting its VA ratings. Second, even similar seeds may have quite different ratings (or an inverse polarity) of valence/arousal to the unseen word, thus reducing prediction performance. To overcome these limitations, this study proposes a community-based weighted graph model that can select seeds which are both similar to and have similar ratings (or the same polarity) with each unseen word to form a community (subgraph) so that its VA ratings can be estimated from such high-quality seeds using a weighted propagation scheme. That is, seeds more similar to unseen words contribute more to the estimation process. Experimental results show that the proposed method yields better prediction performance for both English and Chinese datasets. Jin Wang 0008, Liang-Chih Yu, K. Robert Lai, Xuejie Zhang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | A locally weighted method to improve linear regression for lexical-based valence-arousal predictionabstractText-based sentiment analysis is a growing research field in affective computing, driven by both commercial applications and academic interest. Continuous dimensional representations, such as valence-arousal (VA) space, can represent the affective state more precisely than discrete effective representations. In building dimensional sentiment applications, affective lexicons with valence-arousal ratings are useful resources but are still very rare. Therefore, recent studies have investigated the automatic development of VA lexicons using linear regression techniques. One of the major limitations of linear regression is the under-fitting problem which can cause a poor fit between the algorithm and the training data. To tackle this problem, this study proposes the use of a locally weighted linear regression (LWLR) model to predict the valence-arousal ratings of affective words. The locally weighted method performs a regression around the point of interest using only training data that are "local" to that point, and thus can reduce the impact of noise from unrelated training data. Experimental results show that the proposed method achieved better performance for VA word prediction. Jin Wang 0008, K. Robert Lai, Liang-Chih Yu, Xuejie Zhang 0002 |
ACII | 4 |
| 2015 | A Task-Type-Based Algorithm for the Energy-Aware Profit Maximizing Scheduling Problem in Heterogeneous Computing SystemsabstractIn this paper, we design an efficient algorithm for the energy-aware profit maximizing scheduling problem, where the high performance computing system administrator is to maximize the profit per unit time. The running time of the proposed algorithm is depending on the number of task types, while the running time of the previous algorithm is depending on the number of tasks. Moreover, we prove that the worst-case performance ratio is close to 2, which maybe the best result. Simulation experiments show that the proposed algorithm is more accurate than the previous method. Weidong Li 0002, Xi Liu 0002, Xuejie Zhang 0002, Xiaobo Cai |
CCGRID | 3 |
| 2015 | Enhanced fast compressive tracking based on adaptive measurement matrixabstractRobust object tracking is a challenging task because of factors such as pose variation, illumination changes, abrupt motion and background clutter across the video sequence. With the introduction of the compressive sensing theory, researchers are provided with a new and effective way of real‐time object tracking. In this study, an enhanced fast compressive tracking based on an adaptive measurement matrix is presented, which the authors have named ‘adaptive fast compressive tracking’ (AFCT). The sparsity of the matrix and the number of columns are adaptively determined according to the dimension of the Haar‐like feature. This measurement matrix is fixed once it has been calculated when selecting a tracked rectangle region in the first frame. Unlike most of the existing compressive trackers, the proposed method adopts a different adaptive measurement matrix for a different targeting object. Compared with the fast compressive tracking (FCT), each measurement element contains more information for the original signal. As a result, stable object tracking is achieved by using fewer measurement elements. The proposed AFCT method can run in real time and outperforms FCT on many challenging video sequences in terms of efficiency, accuracy and robustness. Xuejie Zhang 0002 |
IET Comput. Vis. | 3 |
| 2015 | Penalty cost constrained identical parallel machine scheduling problem
Weidong Li 0002, Jianping Li 0007, Xuejie Zhang 0002 |
Theor. Comput. Sci. | 3 |
| 2012 | An Effective Partition Approach for Elastic Application Development on Mobile Cloud Computing
Zhuoran Qin, Jixian Zhang 0003, Xuejie Zhang 0002 |
GPC | 3 |
| 2010 | A Teaching Schema for Multi-core Programming
Xuejie Zhang 0002 |
CSEDU (2) | 2 |
| 2007 | Double Token-Ring and Region-Tree Based Group Communication Mechanism for Mobile Agent
Xuejie Zhang 0002 |
PRIMA | 2 |
| 2006 | Multi Region-Tree Based Dynamic Commission Home Proxy Communication Mechanism for Mobile Agent
Xuejie Zhang 0002 |
PRIMA | 2 |