Shezheng Song

dblp:302/7694 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
16since 2021 · last 2026
0009-0007-9985-7619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging Image as Compressed Visual Prompt and Hierarchical Visual Knowledge for Effective Image Utilization in MLLMs
abstract
Multimodal Large Language Models (MLLMs) integrate text and images for complex reasoning tasks, but efficiently utilizing image remains a challenge due to redundancy and noise. Traditional methods take the entire image features as visual prompt into the MLLMs, leading to excessive visual tokens that disrupt textual information expression. Thus, recent studies treat image features as visual knowledge, storing them in the feed-forward network for retrieval when needed. These methods, completely removing images from the input, may hinder the activation of image-related knowledge. Besides, current visual knowledge focuses on fine-grained details but overlooks the hierarchical process of visual perception. As described in feature integration theory, global structure is first processed before details are integrated. Ignoring this process may lead to a fragmented visual understanding, making it difficult to capture high-level semantic relationships. To overcome these issues, we propose a novel image utilization mechanism in MLLMs. We leverage a compression-based attention mechanism to generate the compressed visual prompt, which not only mitigates the interference of excessively long visual prompts but also preserves crucial visual information necessary for activating knowledge in the MLLM. Furthermore, we extract hierarchical visual features as visual knowledge using wavelet transforms, allowing the model to capture both global structures and fine-grained details. Experiments show that our method achieves state-of-the-art performance.
Shezheng Song, Kangcheng Ding, Shan Zhao 0002, Shasha Li 0001, Xiaopeng Li 0006, Chengyu Wang 0008, Qian Wan 0007, Bin Ji 0002, Jie Yu 0008
AAAI1
2026 EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
abstract
Large Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted parameter modifications, with meta-learning-based model editing (MLME) demonstrating strong effectiveness and efficiency. However, we find that MLME struggles in low-data regimes and incurs high training costs due to the use of KL divergence. To address these issues, we propose $\textbf{E}$fficient $\textbf{M}$ulti-$\textbf{S}$tep $\textbf{Edit (EMSEdit)}$, which leverages multi-step backpropagation (MSBP) to effectively capture gradient-activation mapping patterns within editing samples, performs multi-step edits per sample to enhance editing performance under limited data, and introduces norm-based regularization to preserve unedited knowledge while improving training efficiency. Experiments on two datasets and three LLMs show that EMSEdit consistently outperforms state-of-the-art methods in both sequential and batch editing. Moreover, MSBP can be seamlessly integrated into existing approaches to yield additional performance gains. Further experiments on a multi-hop reasoning editing task demonstrate EMSEdit's robustness in handling complex edits, while ablation studies validate the contribution of each design component. Our code is available at https://github.com/xpq-tech/emsedit.
Xiaopeng Li 0006, Shasha Li 0001, Xi Wang 0018, Shezheng Song, Bin Ji 0002, Shangwen Wang, Jun Ma 0015, Xiaodong Liu 0004, Mina Liu, Jie Yu 0008
WWW4
2026 RICA: Re-ranking with intra-modal and cross-modal alignment for text-based person search
Yu Bai 0022, Wentao Ma 0003, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Chengyu Wang 0008, Qian Wan 0007
Expert Syst. Appl.5
2025 SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
abstract
The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability.
Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Huijun Liu 0003, Bin Ji 0002, Xi Wang 0018, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004
AAAI3
2025 Identifying Knowledge Editing Types in Large Language Models
abstract
Warning: This paper contains examples of toxic text. Knowledge editing has emerged as an efficient technique for updating the knowledge of large language models (LLMs), attracting increasing attention in recent years. However, there is a lack of effective measures to prevent the malicious misuse of this technique, which could lead to harmful edits in LLMs. These malicious modifications could cause LLMs to generate toxic content, misleading users into inappropriate actions. In front of this risk, we introduce a new task, Knowledge Editing Type Identification (KETI), aimed at identifying different types of edits in LLMs, thereby providing timely alerts to users when encountering illicit edits. As part of this task, we propose KETIBench, which includes five types of harmful edits covering the most popular toxic types, as well as one benign factual edit. We develop five classical classification models and three BERT-based models as baseline identifiers for both open-source and closed-source LLMs. Our experimental results, across 92 trials involving four models and three knowledge editing methods, demonstrate that all eight baseline identifiers achieve decent identification performance, highlighting the feasibility of identifying malicious edits in LLMs. Additional analyses reveal that the performance of the identifiers is independent of the reliability of the knowledge editing methods and exhibits cross-domain generalization, enabling the identification of edits from unknown sources. All data and code are available in https://github.com/xpq-tech/KETI.
Xiaopeng Li 0006, Shasha Li 0001, Shangwen Wang, Shezheng Song, Bin Ji 0002, Huijun Liu 0003, Jun Ma 0015, Jie Yu 0008
KDD (2)4
2025 Rethinking Residual Distribution in Locate-then-Edit Model Editing
abstract
Model editing enables targeted updates to the knowledge of large language models (LLMs) with minimal retraining. Among existing approaches, locate-then-edit methods constitute a prominent paradigm: they first identify critical layers, then compute residuals at the final critical layer based on the target edit, and finally apply least-squares-based multi-layer updates via $\textbf{residual distribution}$. While empirically effective, we identify a counterintuitive failure mode: residual distribution, a core mechanism in these methods, introduces weight shift errors that undermine editing precision. Through theoretical and empirical analysis, we show that such errors increase with the distribution distance, batch size, and edit sequence length, ultimately leading to inaccurate or suboptimal edits. To address this, we propose the $\textbf{B}$oundary $\textbf{L}$ayer $\textbf{U}$pdat$\textbf{E (BLUE)}$ strategy to enhance locate-then-edit methods. Sequential batch editing experiments on three LLMs and two datasets demonstrate that BLUE not only delivers an average performance improvement of 35.59\%, significantly advancing the state of the art in model editing, but also enhances the preservation of LLMs' general capabilities. Our code is available at https://github.com/xpq-tech/BLUE.
Xiaopeng Li 0006, Shangwen Wang, Shasha Li 0001, Shezheng Song, Bin Ji 0002, Ma Jun, Jie Yu 0008
NeurIPS4
2025 Psychologically-Aware Retrieval-Augmented Generation for Coherent Role-Playing in LLMs
Pengyang Shao, Shan Zhao 0002, Shezheng Song, Tianwei Yan 0001, Chengyu Wang 0008
PRCV (4)4
2025 How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model
abstract
We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, MLLMs still face challenges in addressing the semantic gap in multimodal data, which may lead to erroneous outputs, posing potential risks to society. Selecting the appropriate modality alignment method is crucial, as improper methods might require more parameters without significant performance improvements. This paper aims to explore modality alignment methods for LLMs and their current capabilities. Implementing effective modality alignment can help LLMs address environmental issues and enhance accessibility. The study surveys existing modality alignment methods for MLLMs, categorizing them into four groups: (1) Multimodal Converter, which transforms data into a format that LLMs can understand; (2) Multimodal Perceiver, which improves how LLMs percieve different types of data; (3) Tool Learning, which leverages external tools to convert data into a common format, usually text; and (4) Data-Driven Method, which teaches LLMs to understand specific data types within datasets.
Shezheng Song, Xiaopeng Li 0006, Shasha Li 0001, Shan Zhao 0002, Jie Yu 0008, Jun Ma 0015, Xiaoguang Mao, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.1
2025 Hierarchical Label-Enhanced Contrastive Learning for Chinese NER
abstract
Recently, character-word lattice structures have achieved promising results for Chinese named entity recognition (NER), reducing word segmentation errors and increasing word boundary information for character sequences. However, constructing the lattice structure is complex and time-consuming, thus these lattice-based models usually suffer from low inference speed. Moreover, the quality of the lexicon affects the accuracy of the NER model. Since noise words can potentially confuse NER, limited coverage of the lexicon can cause lattice-based models to degenerate into partial character-based models. In this article, we propose a hierarchical label-enhanced contrastive learning (HLCL) method for Chinese NER. Instead of relying on the lattice structure, HLCL offers an alternative solution to robustly integrate entity boundary and type information with the help of both labels semantic and contrastive learning. HLCL is empowered by two techniques: 1) sentence-level contrastive learning (SCL) to model global mutual information between two different modalities (e.g., labels and sentences) and 2) token-level contrastive learning (TCL) to close the gap between representations of different characters (e.g., label-enhanced characters and original characters), resulting in local mutual information. With the well-designed contrastive learning scheme and the concise model during inference, HLCL can fully leverage the transferable label semantic and has a superb speed of inference. Experiments on four Chinese NER datasets show that HLCL obtains excellent efficiency as well as performance compared with existing lattice-based approaches.
Chengyu Wang 0008, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Wentao Ma 0003, Kuien Liu, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 FRCL-MNER: A Finer Grained Rank-Based Contrastive Learning Framework for Multimodal NER
abstract
Multimodal named entity recognition (MNER) is an emerging field that aims to automatically detect named entities and classify their categories, utilizing input text and auxiliary resources such as images. While previous studies have leveraged object detectors to preprocess images and fuse textual semantics with corresponding image features, these methods often overlook the potential finer grained information within each modality and may exacerbate error propagation due to predetection. To address these issues, we propose a finer grained rank-based contrastive learning (FRCL) framework for MNER. This framework employs a global-level contrastive learning to align multimodal semantic features and a Top-K rank-based mask strategy to construct positive-negative pairs, thereby learning a finer grained multimodal interaction representation. Experimental results from three well-known social media datasets reveal that our approach surpasses existing strong baselines, and achieves up to a 1.54% improvement on the Twitter2015 dataset. Extensive discussions further confirm the effectiveness of our approach. We will release the source code on https://github.com/augusyan/FRCL.
Tianwei Yan 0001, Shan Zhao 0002, Wentao Ma 0003, Shezheng Song, Chengyu Wang 0008, Zhibo Rao, Shizhao Chen, Zhigang Luo, Xinwang Liu 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 PMET: Precise Model Editing in a Transformer
abstract
Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usually optimize the TL hidden states to memorize target knowledge and use it to update the weights of the FFN in LLMs. However, the information flow of TL hidden states comes from three parts: Multi-Head Self-Attention (MHSA), FFN, and residual connections. Existing methods neglect the fact that the TL hidden states contains information not specifically required for FFN. Consequently, the performance of model editing decreases. To achieve more precise model editing, we analyze hidden states of MHSA and FFN, finding that MHSA encodes certain general knowledge extraction patterns. This implies that MHSA weights do not require updating when new knowledge is introduced. Based on above findings, we introduce PMET, which simultaneously optimizes Transformer Component (TC, namely MHSA and FFN) hidden states, while only using the optimized TC hidden states of FFN to precisely update FFN weights. Our experiments demonstrate that PMET exhibits state-of-the-art performance on both the \textsc{counterfact} and zsRE datasets. Our ablation experiments substantiate the effectiveness of our enhancements, further reinforcing the finding that the MHSA encodes certain general knowledge extraction patterns and indicating its storage of a small amount of factual knowledge. Our code is available at \url{https://github.com/xpq-tech/PMET}.
Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Jun Ma 0015, Jie Yu 0008
AAAI3
2024 A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
abstract
Multimodal Entity Linking (MEL) aims at linking ambiguous mentions with multimodal information to entity in Knowledge Graph (KG) such as Wikipedia, which plays a key role in many applications. However, existing methods suffer from shortcomings, including modality impurity such as noise in raw image and ambiguous textual entity representation, which puts obstacles to MEL. We formulate multimodal entity linking as a neural text matching problem where each multimodal information (text and image) is treated as a query, and the model learns the mapping from each query to the relevant entity from candidate entities. This paper introduces a dual-way enhanced (DWE) framework for MEL: (1) our model refines queries with multimodal data and addresses semantic gaps using cross-modal enhancers between text and image information. Besides, DWE innovatively leverages fine-grained image attributes, including facial characteristic and scene feature, to enhance and refine visual features. (2)By using Wikipedia descriptions, DWE enriches entity semantics and obtains more comprehensive textual representation, which reduces between textual representation and the entities in KG. Extensive experiments on three public benchmarks demonstrate that our method achieves state-of-the-art (SOTA) performance, indicating the superiority of our model. The code is released on https://github.com/season1blue/DWE.
Shezheng Song, Shan Zhao 0002, Chengyu Wang 0008, Tianwei Yan 0001, Shasha Li 0001, Xiaoguang Mao, Meng Wang 0001
AAAI1
2024 DIM: Dynamic Integration of Multimodal Entity Linking with Large Language Model
Shezheng Song, Shasha Li 0001, Jie Yu 0008, Shan Zhao 0002, Xiaopeng Li 0006, Jun Ma 0015, Xiaodong Liu 0004, Xiaoguang Mao
PRCV (5)1
2021 T-Mask: An Active and Accurate Dialogue State Tracking with Token Mask Prediction
abstract
Recent dialogue state tracking (DST) usually treats utterance, system action and ontology equally to estimate the slot types and values. In this way, the expression of slot in utterance is restricted. As the main way to directly express user semantics, utterance should receive further attention and its proportion in semantic expression should be dynamic according to the content. It’s common to recognize the different importance of information in all the DST models. However, most of them pay little attention to the position of slot in utterance. In fact, position and semantics are related due to human grammatical habits and expression habits. Therefore, we propose T-Mask1, a model to actively and accurately learn the token mask position of slot, and we further utilize the learned position information to influence the semantic expression of utterance. We verify the effectiveness of our model on DSTC2 and WoZ2.0. On WoZ2.0, we achieve 90.84 joint goal accuracy and 97.6 turn request accuracy, which is better than most existing models.
Shezheng Song, Dongsong Zhang, Zhen Huang 0006, Yuxing Peng 0001
ICTAI1
2021 SEED: A Cross-Layer Semantic Enhanced SLU Model With Role Context Differentiated Fusion
abstract
The mainstream SLU models, such as SDEN, take the joint training way of slot filling and intent detection because of their correlation and add contextual information to improve the model performance by the contextual vector. Although these models have proved effective, it also brings challenges for slot filling. The slot filling decoder is fed with the deep-layer semantic encoding without alignment information, which will affect the performance of slot filling. The alignment information of the history utterances is attenuated in the context vector because of the repeated fusion process, which is not conducive to the performance improvement of slot filling. In order to solve the above problems, we proposed a novel cross layer semantic enhanced SLU model with role context differentiated fusion, which contains two important improvements: 1) the word embedding information of the current utterance is introduced into the slot filling decoder to strengthen the alignment information based on the mutual attention mechanism; 2) the utterances of different roles are fused in different ways to reserve the alignment information of history utterances in the contextual vector. A large number of experiments were carried out on the standard dataset from SDEN, named KVRET*, and the results verify the effectiveness of our new model. Our model can increase the F1 score of slot filling by more than 7.5% than the existing models.
Dongsong Zhang, Shezheng Song, Zhen Huang 0006, Yuxing Peng 0001
ICTAI3
2021 USET : A network based on Utterance hidden State transfEr for Task-oriented dialogue
abstract
Multi-turn dialogue is challenging because semantic information is not only contained in the current utterance, but also in the dialogue context. In fact, understanding multiturn dialogue is a dynamic process. With the increase of dialogue turn, users' understanding is also changing. In this case, we propose a network based on utterance hidden state transfer for task-oriented dialogue (USET). In our model, we first extract the hidden state of previous utterance as the previous comprehension. Then, this comprehension is passed to the next turn. We take the comprehension as a prior knowledge to understand the semantic information in the dialogue context. Finally, we put the previous comprehension and current utterance together to understand current utterance. In order to realize the transfer of comprehension in dialogue, we propose a continuous sample training method: CST, which takes a multi-turn dialogue as a whole to understand. All the sentences in a dialogue are put into a batch for training. Our method makes use of previous comprehension and achieves the information exchange among dialogue. Experimental results on Stanford Multi-Domain dataset demonstrate that our model is superior to existing models. Code is available at https://github.com/season1blue/USET
Shezheng Song, Dongsong Zhang, Yuxing Peng 0001, Yuan Yuan 0034
IJCNN1