Peichao Lai

dblp:331/1057 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-6936-5687ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Entity injection with contrastive learning encoder for Chinese few-shot natural language inference
Peichao Lai, Feiyang Ye 0002, Yanggeng Fu, Ruiqing Wang
Appl. Intell.1
2026 Cluster-based prototypical contrastive learning for unsupervised sentence embedding
Peichao Lai, Ruiqing Wang, Jiayong Li, Ruixiong Fang, Zhengfeng Zhang, Qingwei Lyu
Eng. Appl. Artif. Intell.1
2026 Quantum-inspired neural networks with stochastic dynamics for multimodal sentiment analysis and sarcasm detection
abstract
Quantum-inspired neural networks have demonstrated strong potential in modeling non-classical phenomena in cognitive tasks, particularly in multimodal sentiment analysis, marking a significant advancement over traditional models. However, existing multimodal quantum-inspired neural networks fall short in fully modeling the multimodal density matrix, typically relying on simplistic neural mappings to represent quantum entanglement. This lack of explicit physical constraints, particularly those governing open quantum system dynamics, limits both the interpretability and performance. To address this limitation, we propose a novel framework grounded in quantum stochastic dynamics, introducing two quantum-inspired neural networks, which model the evolution of multimodal data as Markovian and non-Markovian open quantum systems, respectively. This approach enables the simulation of quantum system evolution to capture rich non-classical interactions between modalities. The resulting entangled multimodal density matrix is then measured through quantum projections to extract high-level features for downstream sentiment analysis and sarcasm detection. Extensive experiments on benchmark bimodal and trimodal datasets demonstrate that our models consistently outperform state-of-the-art traditional baselines, large-scale language models and quantum-inspired neural networks. Ablation studies confirm the critical role of quantum stochastic dynamics in performance gains. Furthermore, we enhance the interpretability by tracking the evolution of the density matrix using von-Neumann entanglement entropy as a quantitative metric, providing deeper insight into the internal mechanisms of the model.
Kehuan Yan, Peichao Lai, Xianghan Zheng, Yi Ren 0001, Tuyatsetseg Badarch, Yiwei Chen 0002
Eng. Appl. Artif. Intell.2
2026 SAS-bench: A fine-grained benchmark for evaluating short answer scoring with large language models
Peichao Lai, Kexuan Zhang, Linyihan Zhang, Feiyang Ye 0002, Jinhao Yan, Yanwei Xu 0004, Conghui He, Wentao Zhang 0001, Bin Cui 0001
Neural Networks1
2026 Improving Low-Resource Short Answer Scoring Through Large Language Model-Based Data Augmentation
abstract
The automated grading of subjective answers is crucial for reducing manual workload and enhancing feedback efficiency in online education, particularly for short answer scoring (SAS). However, in scenarios with limited labeled data, existing methods face challenges in sample diversity and scoring consistency due to limited training data and misaligned label distributions. While data augmentation and transfer learning have been employed to address these issues, rule-based approaches often lack textual variability, and general-domain embeddings struggle to align with domain-specific scoring criteria. To overcome these limitations, we propose the Scoring with Contextual Alignment and Language Enhancement (SCALE) framework, a novel LLM-driven training paradigm that synthesizes diverse responses while preserving scoring consistency. SCALE leverages a knowledge graph-based generation strategy to enhance sample diversity by substituting key phrases with contextually aligned alternatives and employs a style rewrite prompt to introduce linguistic variations. To mitigate label inconsistency, we introduce a polish align prompt that refines synthetic and real samples into a shared semantic subspace, training an annotator model for aligned scoring. Additionally, an entity-aware enhancement mechanism improves comprehension of formulas and quantitative content. Extensive experiments on multilingual and multi-domain datasets demonstrate that SCALE achieves state-of-the-art performance, improving Pearson scores by 4.9%, 2.27%, and 1.34% over BERT, RoBERTa, and ERNIE 3.0, respectively.
Peichao Lai, Kexuan Zhang, Bin Cui 0001
IEEE Trans. Knowl. Data Eng.1
2025 Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning
abstract
Recently, using large language models (LLMs) for data augmentation has led to considerable improvements in unsupervised sentence embedding models. However, existing methods encounter two primary challenges: limited data diversity and high data noise. Current approaches often neglect fine-grained knowledge, such as entities and quantities, leading to insufficient diversity. Besides, unsupervised data frequently lacks discriminative information, and the generated synthetic samples may introduce noise. In this paper, we propose a pipeline-based data augmentation method via LLMs and introduce the Gaussian-decayed gradient-assisted Contrastive Sentence Embedding (GCSE) model to enhance unsupervised sentence embeddings. To tackle the issue of low data diversity, our pipeline utilizes knowledge graphs (KGs) to extract entities and quantities, enabling LLMs to generate more diverse samples. To address high data noise, the GCSE model uses a Gaussian-decayed function to limit the impact of false hard negative samples, enhancing the model’s discriminative capability. Experimental results show that our approach achieves state-of-the-art performance in semantic textual similarity (STS) tasks, using fewer data samples and smaller LLMs, demonstrating its efficiency and robustness across various models.
Peichao Lai, Zhengfeng Zhang, Wentao Zhang 0001, Fangcheng Fu, Bin Cui 0001
ACL (1)1
2025 Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations
abstract
Sequence labeling remains a significant challenge in low-resource, domain-specific scenarios, particularly for character-dense languages.Existing methods primarily focus on enhancing model comprehension and improving data diversity to boost performance.However, these approaches still struggle with inadequate model applicability and semantic distribution biases in domain-specific contexts.To overcome these limitations, we propose a novel framework that combines an LLM-based knowledge enhancement workflow with a span-based Knowledge Fusion for Rich and Efficient Extraction (KnowFREE) model 1 .Our workflow employs explanation prompts to generate precise contextual interpretations of target entities, effectively mitigating semantic biases and enriching the model's contextual understanding.The KnowFREE model further integrates extension label features, enabling efficient nested entity extraction without relying on external knowledge during inference.Experiments on multiple domain-specific sequence labeling datasets demonstrate that our approach achieves stateof-the-art performance, effectively addressing the challenges posed by low-resource settings.
Peichao Lai, Jiaxin Gan, Feiyang Ye 0002, Wentao Zhang 0001, Fangcheng Fu, Bin Cui 0001
EMNLP1
2025 ADSC: LLM-Augmented Dual-Stream Cooperative Learning for Robust Automated Essay Scoring
Kexuan Zhang, Peichao Lai, Jikai Hu
PRICAI2
2025 FE-CFNER: Feature Enhancement-based approach for Chinese Few-shot Named Entity Recognition
Sanhe Yang, Peichao Lai, Ruixiong Fang, Yanggeng Fu, Feiyang Ye 0002
Comput. Speech Lang.2
2025 Span-based sliding window feature extraction and label knowledge for few-shot Named Entity Recognition
Jiaxin Gan, Ruixiong Fang, Peichao Lai
Eng. Appl. Artif. Intell.4
2025 Quantum-inspired multimodal fusion with Lindblad master equation for sentiment analysis
Kehuan Yan, Peichao Lai, Yi Ren 0001, Tuyatsetseg Badarch, Yiwei Chen 0002, Xianghan Zheng
Neurocomputing2
2024 Quantum-inspired Neural Network Based on Stochastic Liouville-von Neumann Equation for Sentiment Classification
abstract
Quantum-inspired models have shown enhanced capabilities in various language tasks, including question answering and sentiment analysis. However, current complex-valued-based models primarily focus on sentence embedding, overlooking the significance of the quantum evolution process and the extra time cost incurred by complex expressions. In this work, we present a novel quantum-inspired neural network, SSS-QNN, which integrates the Stochastic Liouville-von Neumann Equation (SLE) to simulate the evolution process and the complex-valued simple recurrent unit (SRU) to reduce the time cost, offering the model physical meaning, thus enhancing the interpretability. We conduct comprehensive experiments on both sentence-level and document-level sentiment classification datasets. Compared to traditional models, large language models, and quantum-inspired models, SSS-QNN demonstrates competitive performance in accuracy and time cost. Additional ablation tests verify the effectiveness of the proposed modules.
Kehuan Yan, Peichao Lai, Qingwei Lyu
IJCNN2
2024 NCSE: Neighbor Contrastive Learning for Unsupervised Sentence Embeddings
abstract
Unsupervised sentence embedding methods based on contrastive learning have gained attention for effectively representing sentences in natural language processing. Retrieving additional samples via a nearest-neighbor approach can enhance the model’s ability to learn relevant semantics and distinguish sentences. However, previous related research mainly focused on retrieving neighboring samples within a single batch range or global range, which makes the model possibly unable to capture effective semantic information or incurs excessive time cost. Furthermore, previous methods use retrieved neighbor samples as hard negatives. We argue that nearest neighbor samples contain relevant semantic information, and treating them as hard negatives risks losing valuable semantic knowledge. In this work, we introduce Neighbor Contrastive learning for unsupervised Sentence Embeddings(NCSE), which combines contrastive learning with the nearest-neighbor approach. Specifically, we create a candidate set to store sentence embeddings across multiple batches. Retrieving the candidate set can ensure sufficient samples, making it easier for the model to learn relevant semantics. Using retrieved nearest neighbor samples as positives and applying the self-attention mechanism to aggregate the sample and its neighbors encourages the model to learn relevant semantics from multiple neighbors. Experiments on the semantic text similarity task demonstrate our method’s effectiveness in sentence embedding learning.
Zhengfeng Zhang, Peichao Lai, Ruiqing Wang, Feiyang Ye 0002
IJCNN2
2024 Quantum-inspired Language Model with Lindblad Master Equation and Interference Measurement for Sentiment Analysis
abstract
Kehuan Yan, Peichao Lai, Yilei Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kehuan Yan, Peichao Lai
NAACL-HLT2
2024 Span-Based Chinese Few-Shot NER with Contrastive and Prompt Learning
Feiyang Ye 0002, Peichao Lai, Sanhe Yang, Zhengfeng Zhang
NLPCC (2)2
2024 M-Sim: Multi-level Semantic Inference Model for Chinese short answer scoring in low-resource scenarios
Peichao Lai, Feiyang Ye 0002, Yanggeng Fu
Comput. Speech Lang.1
2024 CogNLG: Cognitive graph for KG-to-text generation
abstract
Abstract Knowledge graph (KG) has been fully considered in natural language generation (NLG) tasks. A KG can help models generate controllable text and achieve better performance. However, most existing related approaches still lack explainability and scalability in large‐scale knowledge reasoning. In this work, we propose a novel CogNLG framework for KG‐to‐text generation tasks. Our CogNLG is implemented based on the dual‐process theory in cognitive science. It consists of two systems: one system acts as the analytic system for knowledge extraction, and another is the perceptual system for text generation by using existing knowledge. During text generation, CogNLG provides a visible and explainable reasoning path. Our framework shows excellent performance on all datasets and achieves a BLEU score of 36.7, which increases by 6.7 compared to the best competitor.
Peichao Lai, Feiyang Ye 0002, Yanggeng Fu, Victor Chang 0001
Expert Syst. J. Knowl. Eng.1
2022 PCBERT: Parent and Child BERT for Chinese Few-shot NER
abstract
Achieving good performance on few-shot or zero-shot datasets has been a long-term challenge for NER. The conventional semantic transfer approaches on NER will decrease model performance when the semantic distribution is quite different, especially in Chinese few-shot NER. Recently, prompt-tuning has been thoroughly considered for low-resource tasks. But there is no effective prompt-tuning approach for Chinese few-shot NER. In this work, we propose a prompt-based Parent and Child BERT (PCBERT) for Chinese few-shot NER. To train an annotating model on high-resource datasets and then discover more implicit labels on low-resource datasets. We further design a label extension strategy to achieve label transferring from high-resource datasets. We evaluated our model on Weibo and the other three sampling Chinese NER datasets, and the experimental result demonstrates our approach’s effectiveness in few-shot learning.
Peichao Lai, Feiyang Ye 0002, Yanggeng Fu
COLING1
2022 Chinese Medical Named Entity Recognition Using External Knowledge
Peichao Lai, Feiyang Ye 0002, Ruixiong Fang, Ruiqing Wang, Jiayong Li
PRICAI (2)2