Wei Song 0010

dblp:62/1539-10 · DBLP profile ↗
← Back
27ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0003-2334-3623ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 CPT-Agent: A Cognitive Process Theory-driven Framework for Student Simulation in Writing Development
abstract
Yuhan Chen, Zizhuo Shen, Miaomiao Cheng, Xu Han, Jiefu Gong, Shijin Wang, Wei Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zizhuo Shen, MiaoMiao Cheng, Xu Han 0021, Jiefu Gong, Shijin Wang 0001, Wei Song 0010
ACL (1)7
2026 Multimodal sentiment analysis with query-based distillation and asymmetric fusion
Wei Song 0010
Neurocomputing4
2025 IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory
abstract
Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks. However, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance and cost. While powerful models deliver better results, they come at a high cost, whereas smaller models are more cost-effective but less capable. To address this trade-off, we propose IRT-Router, a multi-LLM routing framework that efficiently routes user queries to the most suitable LLM. Inspired by Item Response Theory (IRT), a psychological measurement methodology, IRT-Router explicitly models the relationship between LLM capabilities and user query attributes. This not only enables accurate prediction of response performance but also provides interpretable insights, such as LLM abilities and query difficulty. Additionally, we design an online query warm-up technique based on semantic similarity, further enhancing the online generalization capability of IRT-Router. Extensive experiments on 20 LLMs and 12 datasets demonstrate that IRT-Router outperforms most baseline methods in terms of effectiveness and interpretability. Its superior performance in cold-start scenarios further confirms the reliability and practicality of IRT-Router in real-world applications. Code is available at https://github.com/Mercidaiha/IRT-Router.
Wei Song 0010, Zhenya Huang, Weibo Gao, Bihan Xu, Guanhao Zhao, Fei Wang 0063, Runze Wu 0001
ACL (1)1
2025 Divergence-enhanced Knowledge-guided Context Optimization for Visual-Language Prompt Tuning
abstract
Prompt tuning vision-language models like CLIP has shown great potential in learning transferable representations for various downstream tasks. The main issue is how to mitigate the over-fitting problem on downstream tasks with limited training samples. While knowledge-guided context optimization has been proposed by constructing consistency constraints to handle catastrophic forgetting in the pre-trained backbone, it also introduces a bias toward pre-training. This paper proposes a novel and simple Divergence-enhanced Knowledge-guided Prompt Tuning (DeKg) method to address this issue. The key insight is that the bias toward pre-training can be alleviated by encouraging the independence between the learnable and the crafted prompt. Specifically, DeKg employs the Hilbert-Schmidt Independence Criterion (HSIC) to regularize the learnable prompts, thereby reducing their dependence on prior general knowledge, and enabling divergence induced by target knowledge. Comprehensive evaluations demonstrate that DeKg serves as a plug-and-play module can seamlessly integrate with existing knowledge-guided context optimization methods and achieves superior performance in three challenging benchmarks. We make our code available at https://github.com/cnunlp/DeKg.
Yilun Li, MiaoMiao Cheng, Xu Han 0021, Wei Song 0010
ICLR4
2025 Advancing Visible-Infrared Person Re-Identification: Synergizing Visual-Textual Reasoning and Cross-Modal Feature Alignment
abstract
Visible-infrared person re-identification (VI-ReID) is a critical cross-modality fine-grained classification task with significant implications for public safety and security applications. Existing VI-ReID methods primarily focus on extracting modality-invariant features for person retrieval. However, due to the inherent lack of texture information in infrared images, these modality-invariant features tend to emphasize global contexts. Consequently, individuals with similar silhouettes are often misidentified, posing potential risks to security systems and forensic investigations. To address this problem, this paper innovatively introduces natural language descriptions to learn the global-local contexts for VI-ReID. Specifically, we design a framework that jointly optimizes visible-infrared alignment plus (VIAP) and visual-textual reasoning (VTR), and introduces local-global joint measure (LJM) to enhance the metric, while proposing a human-LLM collaborative approach to incorporate textual descriptions into existing cross-modal person re-identification datasets. VIAP achieves cross-modal alignment between RGB and IR. It can explicitly utilize designed frequency-aware modality alignment and relationship-reinforced fusion to explore the potential of local cues in global features and modality-invariant information. VTR proposes pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. LJM proposes introducing local feature distances into the measure stage metric to enhance the relevance of matching using fine-grained information. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches. The dataset is publicly available athttps://github.com/qyx596/vireid-caption.
Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002
IEEE Trans. Inf. Forensics Secur.3
2024 MindMap: Constructing Evidence Chains for Multi-Step Reasoning in Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, they still face significant challenges in automated reasoning, particularly in scenarios involving multi-step reasoning. In this paper, we focus on the logical reasoning problem. The main task is to answer a question based on a set of available facts and rules. A lot of work has focused on guiding LLMs to think logically by generating reasoning paths, ignoring the structure among available facts. In this paper, we propose a simple approach MindMap by introducing evidence chains for supporting reasoning. An evidence chain refers to a set of facts that involve the same subject. In this way, we can organize related facts together to avoid missing important information. MindMap can be integrated with existing reasoning framework, such as Chain-of-Thought (CoT) and Selection-Inference (SI), by letting the model select relevant evidence chains instead of independent facts. The experimental results on the bAbI and ProofWriterOWA datasets demonstrate the effectiveness of MindMap.It can significantly improve CoT and SI, especially in multi-step reasoning tasks.
Yangyu Wu 0001, Xu Han 0021, Wei Song 0010, MiaoMiao Cheng
AAAI3
2024 High-Order Semantic Alignment for Unsupervised Fine-Grained Image-Text Retrieval
abstract
Cross-modal retrieval is an important yet challenging task due to the semantic discrepancy between visual content and language. To measure the correlation between images and text, most existing research mainly focuses on learning global or local correspondence, failing to explore fine-grained local-global alignment. To infer more accurate similarity scores, we introduce a novel High Order Semantic Alignment (HOSA) model that can provide complementary and comprehensive semantic clues. Specifically, to jointly learn global and local alignment and emphasize local-global interaction, we employ tensor-product (t-product) operation to reconstruct one modal’s representation based on another modal’s information in a common semantic space. Such a cross-modal reconstruction strategy would significantly enhance inter-modal correlation learning in a fine-grained manner. Extensive experiments on two benchmark datasets validate that our model significantly outperforms several state-of-the-art baselines, especially in retrieving the most relevant results.
MiaoMiao Cheng, Xu Han 0021, Wei Song 0010
LREC/COLING4
2024 Optimizing Chinese Lexical Simplification Across Word Types: A Hybrid Approach
abstract
This paper addresses the task of Chinese Lexical Simplification (CLS).A key challenge in CLS is the scarcity of data resources.We begin by evaluating the performance of various language models at different scales in unsupervised and few-shot settings, finding that their effectiveness is sensitive to word types.Expensive large language models (LLMs), such as GPT-4, outperform small models in simplifying complex content words and Chinese idioms from the dictionary.To take advantage of this, we propose an automatic knowledge distillation framework called PivotKD for generating training data to fine-tune small models.In addition, all models face difficulties with out-ofdictionary (OOD) words such as internet slang.To address this, we implement a retrieval-based interpretation augmentation (RIA) strategy, injecting word interpretations from external resources into the context.Experimental results demonstrate that fine-tuned small models outperform GPT-4 in simplifying complex content words and Chinese idioms.Additionally, the RIA strategy enhances the performance of most models, particularly in handling OOD words.Our findings suggest that a hybrid approach could optimize CLS performance while managing inference costs.This would involve configuring choices such as model scale, linguistic resources, and the use of RIA based on specific word types to strike an ideal balance.
Zihao Xiao 0004, Jiefu Gong, Shijin Wang 0001, Wei Song 0010
EMNLP4
2024 Joint Visual-Textual Reasoning and Visible-Infrared Modality Alignment for Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a cross-modality fine-grained classification task. Existing approaches for VI-ReID mainly explore modality-invariant features for person retrieval. However, modality-invariant features pay more attention to global contexts, due to the lack of texture information in infrared images. This leads to a person with similar silhouette often being misidentified. Targeting this problem, this paper innovatively introduces natural language specification to learn global-local contexts for VI-ReID. Specifically, our framework jointly optimizes visible-infrared alignment (VIA) and visual-textual reasoning (VTR). VIA achieves cross-modal between RGB and IR. It can explicitly utilize designed modality-guided alignment and relationship-reinforced fusion to explore the potential of local cues in global features. VTR proposes the pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches.
Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002
ICME3
2024 Towards Accurate and Fair Cognitive Diagnosis via Monotonic Data Augmentation
abstract
Intelligent education stands as a prominent application of machine learning. Within this domain, cognitive diagnosis (CD) is a key research focus that aims to diagnose students' proficiency levels in specific knowledge concepts. As a crucial task within the field of education, cognitive diagnosis encompasses two fundamental requirements: accuracy and fairness. Existing studies have achieved significant success by primarily utilizing observed historical logs of student-exercise interactions. However, real-world scenarios often present a challenge, where a substantial number of students engage with a limited number of exercises. This data sparsity issue can lead to both inaccurate and unfair diagnoses. To this end, we introduce a monotonic data augmentation framework, CMCD, to tackle the data sparsity issue and thereby achieve accurate and fair CD results. Specifically, CMCD integrates the monotonicity assumption, a fundamental educational principle in CD, to establish two constraints for data augmentation. These constraints are general and can be applied to the majority of CD backbones. Furthermore, we provide theoretical analysis to guarantee the accuracy and convergence speed of CMCD. Finally, extensive experiments on real-world datasets showcase the efficacy of our framework in addressing the data sparsity issue with accurate and fair CD results.
Zheng Zhang 0048, Wei Song 0010, Qi Liu 0003, Qingyang Mao, Yiyan Wang, Weibo Gao, Zhenya Huang, Shijin Wang 0001, Enhong Chen
NeurIPS2
2024 Discriminative explicit instance selection for implicit discourse relation classification
Wei Song 0010, Hongfei Han, Xu Han 0021, MiaoMiao Cheng, Jiefu Gong, Shijin Wang 0001, Ting Liu 0001
Frontiers Comput. Sci.1
2021 Verb Metaphor Detection via Contextual Relation Learning
abstract
Wei Song, Shuhui Zhou, Ruiji Fu, Ting Liu, Lizhen Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Song 0010, Shuhui Zhou, Ruiji Fu, Ting Liu 0001, Lizhen Liu
ACL/IJCNLP (1)1
2021 A Knowledge Graph Embedding Approach for Metaphor Processing
abstract
Metaphor is a figure of speech that describes one thing (a target) by mentioning another thing (a source) in a way that is not literally true. Metaphor understanding is an interesting but challenging problem in natural language processing. This paper presents a novel method for metaphor processing based on knowledge graph (KG) embedding. Conceptually, we abstract the structure of a metaphor as an attribute-dependent relation between the target and the source. Each specific metaphor can be represented as a metaphor triple (target, attribute, source). Therefore, we can model metaphor triples just like modeling fact triples in a KG and exploit KG embedding techniques to learn better representations of concepts, attributes and concept relations. In this way, metaphor interpretation and generation could be seen as KG completion, while metaphor detection could be viewed as a representation learning enhanced concept pair classification problem. Technically, we build a Chinese metaphor KG in the form of metaphor triples based on simile recognition, and also extract concept-attribute collocations to help describe concepts and measure concept relations. We extend the translation-based and the rotation-based KG embedding models to jointly optimize metaphor KG embedding and concept-attribute collocation embedding. Experimental results demonstrate the effectiveness of our method. Simile recognition is feasible for building the metaphor triple resource. The proposed models improve the performance on metaphor interpretation and generation, and the learned representations also benefit nominal metaphor detection compared with strong baselines.
Wei Song 0010, Jingjin Guo, Ruiji Fu, Ting Liu 0001, Lizhen Liu
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Discourse Self-Attention for Discourse Element Identification in Argumentative Student Essays
abstract
This paper proposes to adapt self-attention to discourse level for modeling discourse elements in argumentative student essays.Specifically, we focus on two issues.First, we propose structural sentence positional encodings to explicitly represent sentence positions.Second, we propose to use inter-sentence attentions to capture sentence interactions and enhance sentence representation.We conduct experiments on two datasets: a Chinese dataset and an English dataset.We find that (i) sentence positional encodings can lead to a large improvement for identifying discourse elements; (ii) a structural relative positional encoding of sentences shows to be most effective; (iii) inter-sentence attention vectors are useful as a kind of sentence representation for identifying discourse elements.
Wei Song 0010, Ziyao Song, Ruiji Fu, Lizhen Liu, MiaoMiao Cheng, Ting Liu 0001
EMNLP (1)1
2020 Multi-Stage Pre-training for Automated Chinese Essay Scoring
abstract
This paper proposes a pre-training based automated Chinese essay scoring method.The method involves three components: weakly supervised pre-training, supervised crossprompt fine-tuning and supervised targetprompt fine-tuning.An essay scorer is first pretrained on a large essay dataset covering diverse topics and with coarse ratings, i.e., good and poor, which are used as a kind of weak supervision.The pre-trained essay scorer would be further fine-tuned on previously rated essays from existing prompts, which have the same score range with the target prompt and provide extra supervision.At last, the scorer is fine-tuned on the target-prompt training data.The evaluation on four prompts shows that this method can improve a state-of-the-art neural essay scorer in terms of effectiveness and domain adaptation ability, while in-depth analysis also reveals its limitations.
Wei Song 0010, Ruiji Fu, Lizhen Liu, Ting Liu 0001, MiaoMiao Cheng
EMNLP (1)1
2020 Hierarchical Multi-task Learning for Organization Evaluation of Argumentative Student Essays
abstract
Organization evaluation is an important dimension of automated essay scoring. This paper focuses on discourse element (i.e., functions of sentences and paragraphs) based organization evaluation. Existing approaches mostly separate discourse element identification and organization evaluation. In contrast, we propose a neural hierarchical multi-task learning approach for jointly optimizing sentence and paragraph level discourse element identification and organization evaluation. We represent the organization as a grid to simulate the visual layout of an essay and integrate discourse elements at multiple linguistic levels. Experimental results show that the multi-task learning based organization evaluation can achieve significant improvements compared with existing work and pipeline baselines. Multiple level discourse element identification also benefits from multi-task learning through mutual enhancement.
Wei Song 0010, Ziyao Song, Lizhen Liu, Ruiji Fu
IJCAI1
2018 Exploiting Syntactic Structures for Humor Recognition
abstract
Humor recognition is an interesting and challenging task in natural language processing. This paper proposes to exploit syntactic structure features to enhance humor recognition. Our method achieves significant improvements compared with humor theory driven baselines. We found that some syntactic structure features consistently correlate with humor, which indicate interesting linguistic phenomena. Both the experimental results and the analysis demonstrate that humor can be viewed as a kind of style and content independent syntactic structures can help identify humor and have good interpretability.
Lizhen Liu, Donghai Zhang, Wei Song 0010
COLING3
2018 Neural Multitask Learning for Simile Recognition
abstract
Simile is a special type of metaphor, where comparators such as like and as are used to compare two objects.Simile recognition is to recognize simile sentences and extract simile components, i.e., the tenor and the vehicle.This paper presents a study of simile recognition in Chinese.We construct an annotated corpus for this research, which consists of 11.3k sentences that contain a comparator.We propose a neural network framework for jointly optimizing three tasks: simile sentence classification, simile component extraction and language modeling.The experimental results show that the neural network based approaches can outperform all rule-based and feature-based baselines.Both simile sentence classification and simile component extraction can benefit from multitask learning.The former can be solved very well, while the latter is more difficult.
Lizhen Liu, Wei Song 0010, Ruiji Fu, Ting Liu 0001
EMNLP3
2018 Semantic composition of distributed representations for query subtopic mining
abstract
Inferring query intent is significant in information retrieval tasks. Query subtopic mining aims to find possible subtopics for a given query to represent potential intents. Subtopic mining is challenging due to the nature of short queries. Learning distributed representations or sequences of words has been developed recently and quickly, making great impacts on many fields. It is still not clear whether distributed representations are effective in alleviating the challenges of query subtopic mining. In this paper, we exploit and compare the main semantic composition of distributed representations for query subtopic mining. Specifically, we focus on two types of distributed representations: paragraph vector which represents word sequences with an arbitrary length directly, and word vector composition. We thoroughly investigate the impacts of semantic composition strategies and the types of data for learning distributed representations. Experiments were conducted on a public dataset offered by the National Institute of Informatics Testbeds and Community for Information Access Research. The empirical results show that distributed semantic representations can achieve outstanding performance for query subtopic mining, compared with traditional semantic representations. More insights are reported as well.
Wei Song 0010, Lizhen Liu, Hanshi Wang
Frontiers Inf. Technol. Electron. Eng.1
2017 Discourse Mode Identification in Essays
abstract
Discourse modes play an important role in writing composition and evaluation.This paper presents a study on the manual and automatic identification of narration, exposition, description, argument and emotion expressing sentences in narrative essays.We annotate a corpus to study the characteristics of discourse modes and describe a neural sequence labeling model for identification.Evaluation results show that discourse modes can be identified automatically with an average F1-score of 0.7.We further demonstrate that discourse modes can be used as features that improve automatic essay scoring (AES).The impacts of discourse modes for AES are also discussed.
Wei Song 0010, Ruiji Fu, Lizhen Liu, Ting Liu 0001
ACL (1)1
2016 Anecdote Recognition and Recommendation
abstract
We introduce a novel task Anecdote Recognition and Recommendation. An anecdote is a story with a point revealing account of an individual person. Recommending proper anecdotes can be used as evidence to support argumentative writing or as a clue for further reading. We represent an anecdote as a structured tuple — < person, story, implication >. Anecdote recognition runs on archived argumentative essays. We extract narratives containing events of a person as the anecdote story. More importantly, we uncover the anecdote implication, which reveals the meaning and topic of an anecdote. Our approach depends on discourse role identification. Discourse roles such as thesis, main ideas and support help us locate stories and their implications in essays. The experiments show that informative and interpretable anecdotes can be recognized. These anecdotes are used for anecdote recommendation. The anecdote recommender can recommend proper anecdotes in response to given topics. The anecdote implication contributes most for bridging user interested topics and relevant anecdotes.
Wei Song 0010, Ruiji Fu, Lizhen Liu, Hanshi Wang, Ting Liu 0001
COLING1
2016 Learning to Identify Sentence Parallelism in Student Essays
abstract
Parallelism is an important rhetorical device. We propose a machine learning approach for automated sentence parallelism identification in student essays. We build an essay dataset with sentence level parallelism annotated. We derive features by combining generalized word alignment strategies and the alignment measures between word sequences. The experimental results show that sentence parallelism can be effectively identified with a F1 score of 82% at pair-wise level and 72% at parallelism chunk level. Based on this approach, we automatically identify sentence parallelism in more than 2000 student essays and study the correlation between the use of sentence parallelism and the types and quality of essays.
Wei Song 0010, Ruiji Fu, Lizhen Liu, Hanshi Wang, Ting Liu 0001
COLING1
2016 Document representation based on semantic smoothed topic model
abstract
The goal of document representation is to capture certain feature of the document. Many existing document representation methods are based on bag-of-words and ignore semantic relevance between words in the document. There we proposed a semantic smoothed topic model to represent document. It takes semantic similarity into consideration for topic of document. We conducted two experiments utilizing this method for text classification and information retrieval task. The experimental results suggest that our method is useful for capturing the semantic of text to alleviating polysemy and synonyms problem and data sparseness problem.
Wei Song 0010, Lizhen Liu, Hanshi Wang
SNPD2
2015 Discourse Element Identification in Student Essays based on Global and Local Cohesion
abstract
We present a method of using cohesion to improve discourse element identification for sentences in student essays.New features for each sentence are derived by considering its relations to global and local cohesion, which are created by means of cohesive resources and subtopic coverage.In our experiments, we obtain significant improvements on identifying all discourse elements, especially of +5% F 1 score on thesis and main idea.The analysis shows that global cohesion can better capture thesis statements.
Wei Song 0010, Ruiji Fu, Lizhen Liu, Ting Liu 0001
EMNLP1
2015 Exploiting Collective Hidden Structures in Webpage Titles for Open Domain Entity Extraction
abstract
We present a novel method for open domain named entity extraction by exploiting the collective hidden structures in webpage titles. Our method uncovers the hidden textual structures shared by sets of webpage titles based on generalized URL patterns and a multiple sequence alignment technique. The highlights of our method include: 1) The boundaries of entities can be identified automatically in a collective way without any manually designed pattern, seed or class name. 2) The connections between entities are also discovered naturally based on the hidden structures, which makes it easy to incorporate distant or weak supervision. The experiments show that our method can harvest large scale of open domain entities with high precision. A large ratio of the extracted entities are long-tailed and complex and cover diverse topics. Given the extracted entities and their connections, we further show the effectiveness of our method in a weakly supervised setting. Our method can produce better domain specific entities in both precision and recall compared with the state-of-the-art approaches.
Wei Song 0010, Hua Wu 0003, Haifeng Wang 0001, Lizhen Liu, Hanshi Wang
WWW1
2012 Multi-aspect query summarization by composite query
abstract
Conventional search engines usually return a ranked list of web pages in response to a query. Users have to visit several pages to locate the relevant parts. A promising future search scenario should involve: (1) understanding user intents; (2) providing relevant information directly to satisfy searchers' needs, as opposed to relevant pages. In this paper, we present a search paradigm to summarize a query's information from different aspects. Query aspects could be aligned to user intents. The generated summaries for query aspects are expected to be both specific and informative, so that users can easily and quickly find relevant information. Specifically, we use a Composite Query for Summarization" method, where a set of component queries are used for providing additional information for the original query. The system leverages the search engine to proactively gather information by submitting multiple component queries according to the original query and its aspects. In this way, we could get more relevant information for each query aspect and roughly classify information. By comparative mining the search results of different component queries, it is able to identify query (dependent) aspect words, which help to generate more specific and informative summaries. The experimental results on two data sets, Wikipedia and TREC ClueWeb2009, are encouraging. Our method outperforms two baseline methods on generating informative summaries.
Wei Song 0010, Zhiheng Xu, Ting Liu 0001, Sheng Li 0003, Ji-Rong Wen
SIGIR1
2011 Query term ranking based on search results overlap
abstract
In this paper, we propose a method to rank and assign weights to query terms according to their impact on the topic of the query. We use Search Result Overlap Ratio (SROR) to quantify the overlap of the search results of the full query and a shorten query after removing one term. Intuitively, if the overlap is small, it indicates a big topic shift and the removed term should be discriminative and important. The SROR could be used for measuring query term importance with a search engine automatically. By this way, learning based models could be trained based on a large number of automatically labeled instances and make predictions for future queries efficiently.
Wei Song 0010, Yu Zhang 0030, Yubin Xie, Ting Liu 0001, Sheng Li 0003
SIGIR1