Peijie Jiang

dblp:28/11328 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 26% Information extraction and text analysis · 20% Trustworthy machine learning · 14%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis · ACL (1) 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models · ICML 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios · EMNLP 2025
Natural language and speech › Language models and text generation
watermarking
0.912025
VLA-Mark: A cross modal watermark for large vision-language alignment models · EMNLP 2025
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
cross-lingual dependency parsing
0.712023
Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency Transfer · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.712023
Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency Transfer · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.712023
Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency Transfer · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Machine learning › Representation and self-supervised learning › pre-training › visual pre-training
boundary-aware pre-training
0.612022
Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling · EMNLP 2022
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.612022
Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling · EMNLP 2022
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.512021
A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging · EMNLP (1) 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation › domain adaptive classification
fine-grained domain adaptation
0.512021
A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.512021
A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
word segmentation
0.512021
A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
feedforward neural network
0.312025
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models · ICML 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.312025
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios · EMNLP 2025
Natural language and speech › Information extraction and text analysis
sequence labeling
0.212022
Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

memory-space visual retracing · 0.9key-value memory · 0.9empirical evaluation · 0.9attribution analysis · 0.9ablation study · 0.9parameter generation network · 0.7curriculum learning · 0.7unsupervised boundary information · 0.6pre-training · 0.6domain-mixed representation learning · 0.5
YearPublicationVenuePosition
2025 Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
abstract
The interpretability of Mixture-of-Experts (MoE) models, especially those with heterogeneous designs, remains underexplored.Existing attribution methods for dense models fail to capture dynamic routing-expert interactions in sparse MoE architectures.To address this issue, we propose a cross-level attribution algorithm to analyze sparse MoE architectures (Qwen 1.5-MoE, OLMoE, Mixtral-8x7B) against dense models (Qwen 1.5-7B, Llama-7B, Mistral-7B).Results show MoE models achieve 31% higher per-layer efficiency via a "mid-activation, lateamplification" pattern: early layers screen experts, while late layers refine knowledge collaboratively.Ablation studies reveal a "basicrefinement" framework-shared experts handle general tasks (entity recognition), while routed experts specialize in domain-specific processing (geographic attributes).Semanticdriven routing is evidenced by strong correlations between attention heads and experts (r = 0.68), enabling task-aware coordination.Notably, architectural depth dictates robustness: deep Qwen 1.5-MoE mitigates expert failures (e.g., 43% MRR drop in geographic tasks when blocking top-10 experts) through shared expert redundancy, whereas shallow OLMoE suffers severe degradation (76% drop).Task sensitivity further guides design: core-sensitive tasks (geography) require concentrated expertise, while distributed-tolerant tasks (object attributes) leverage broader participation.These insights advance MoE interpretability, offering principles to balance efficiency, specialization, and robustness.
Junzhuo Li, Xiuze Zhou, Peijie Jiang, Xuming Hu
ACL (1)4
2025 Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
abstract
Yunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou, Yanggan Gu, Jungang Li, Jingyu Wang, Peijie Jiang, Aiwei Liu, Jia Liu, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yunkai Dang, Mengxi Gao, Xin Zou 0001, Yanggan Gu, Jungang Li, Peijie Jiang, Aiwei Liu, Xuming Hu
EMNLP8
2025 VLA-Mark: A cross modal watermark for large vision-language alignment models
abstract
Shuliang Liu, Zheng Qi, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zheng Qi, Jesse Jiaxi Xu, He Geng, Aiwei Liu, Peijie Jiang, Yik-Cheung Tam, Xuming Hu
EMNLP8
2025 Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
abstract
Despite their impressive capabilities, Multimodal Large Language Models (MLLMs) are prone to hallucinations, i.e., the generated content that is nonsensical or unfaithful to input sources. Unlike in LLMs, hallucinations in MLLMs often stem from the sensitivity of text decoder to visual tokens, leading to a phenomenon akin to "amnesia" about visual information. To address this issue, we propose MemVR, a novel decoding paradigm inspired by common cognition: when the memory of an image seen the moment before is forgotten, people will look at it again for factual answers. Following this principle, we treat visual tokens as supplementary evidence, re-injecting them into the MLLM through Feed Forward Network (FFN) as “key-value memory” at the middle trigger layer. This look-twice mechanism occurs when the model exhibits high uncertainty during inference, effectively enhancing factual alignment. Comprehensive experimental evaluations demonstrate that MemVR significantly mitigates hallucination across various MLLMs and excels in general benchmarks without incurring additional time overhead.
Xin Zou 0001, Yuanhuiyi Lyu, Kening Zheng, Sirui Huang, Junkai Chen, Peijie Jiang, Chang Tang, Xuming Hu
ICML8
2023 Fine-Grained Domain Adaptation for Chinese Syntactic Processing
abstract
Syntactic processing is fundamental to natural language processing. It provides rich and comprehensive syntax information in sentences that could be potentially beneficial for downstream tasks. Recently, pretrained language models have shown great success in Chinese syntactic processing, which typically involves word segmentation, POS tagging, and dependency parsing. However, the on-going research never ends since performance would be degraded drastically when tested on a highly-discrepant domain. This problem is widely accepted as domain adaptation, where the test domain differs from the training domain in supervised learning. Self-training is one promising solution for it, and straightforward source-to-target adaptation has already shown remarkable effectiveness in previous work. While this strategy ignores the fact that sentences of the target domain sentences may have very different gaps from the source training domain. More specifically, sentences with large gaps might fail by direct self-training adaptation. To this end, we propose fine-grained domain adaptation for Chinese syntactic processing in this work, aiming to model the gaps between the source and the target domains accurately and progressively. The key idea is to divide the target domain into fine-grained subdomains by using a specified domain distance metric, and then perform gradual self-training on the subdomains. We further offer an intuitive theoretical illustration based on the theory of Kumar et al. (2020) approximately. In addition, a novel representation learning framework is proposed to encode fine-grained subdomains effectively, aiming to utilize the above idea fully. Experimental results on benchmark datasets show that our method can achieve significant improvements over a variety of baselines.
Meishan Zhang, Peiming Guo, Peijie Jiang, Dingkun Long, Yueheng Sun, Pengjun Xie, Min Zhang 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency Transfer
abstract
Unsupervised cross-lingual transfer has been shown great potentials for dependency parsing of the low-resource languages when there is no annotated treebank available. Recently, the self-training method has received increasing interests because of its state-of-the-art performance in this scenario. In this work, we advance the method further by coupling it with curriculum learning, which guides the self-training in an easy-to-hard manner. Concretely, we present a novel metric to measure the instance difficulty of a dependency parser which is trained mainly on a Treebank from a resource-rich source language. By using the metric, we divide a low-resource target language into several fine-grained sub-languages by their difficulties, and then apply iterative-self-training progressively on these sub-languages. To fully explore the auto-parsed training corpus from sub-languages, we exploit an improved parameter generation network to model the sub-languages for better representation learning. Experimental results show that our final curriculum-style self-training can outperform a range of strong baselines, leading to new state-of-the-art results on unsupervised cross-lingual dependency parsing. We also conduct detailed experimental analyses to examine the proposed approach in depth for comprehensive understandings.
Peiming Guo, Shen Huang, Peijie Jiang, Yueheng Sun, Meishan Zhang, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling
abstract
Boundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition.Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information.However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored.In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT).We apply BABERT for feature induction of Chinese sequence labeling tasks.Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets.In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information.
Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang 0005
EMNLP1
2021 A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging
abstract
Domain adaption for word segmentation and POS tagging is a challenging problem for Chinese lexical processing.Self-training is one promising solution for it, which struggles to construct a set of high-quality pseudo training instances for the target domain.Previous work usually assumes a universal sourceto-target adaption to collect such pseudo corpus, ignoring the different gaps from the target sentences to the source domain.In this work, we start from joint word segmentation and POS tagging, presenting a fine-grained domain adaption method to model the gaps accurately.We measure the gaps by one simple and intuitive metric, and adopt it to develop a pseudo target domain corpus based on finegrained subdomains incrementally.A novel domain-mixed representation learning model is proposed accordingly to encode the multiple subdomains effectively.The whole process is performed progressively for both corpus construction and model training.Experimental results on a benchmark dataset show that our method can gain significant improvements over a vary of baselines.Extensive analyses are performed to show the advantages of our final domain adaption model as well.
Peijie Jiang, Dingkun Long, Yueheng Sun, Meishan Zhang, Pengjun Xie
EMNLP (1)1