Shisong Chen

dblp:161/4279 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0005-0368-6866ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Code LLMs Still Fall Short of Top Programmers: Evaluating Algorithmic Code Generation Through Computational Thinking
abstract
Evaluating the coding capabilities of models through algorithmic code generation is challenging, as it requires deep problem understanding and complex algorithm design. Current benchmarks suffer from a narrow focus on final execution results (such as pass@k), neglecting the crucial reasoning and problem-solving processes inherent in code generation. To address this limitation, we introduce a multi-phase algorithmic code generation benchmark, MUPA, structured around human computational thinking. MUPA dissects the evaluation into four distinct phases: example understanding, algorithm selection, solution description, and code generation. This framework facilitates a comprehensive assessment by providing insights into the model's intermediate problem-solving steps, rather than just the final code. We manually curated 197 high-quality competitive programming problems from Codeforces. Utilizing an LLM-as-a-judge paradigm with specialized prompts, our rigorous evaluation of several existing code generation LLMs reveals significant across-the-board challenges. Notably, we establish a positive correlation, indicating that proficiency in an earlier phase directly impacts performance in subsequent phases, underscoring the interdependency of these algorithmic skills. The benchmark is publicly available at https://github.com/cheniison/MUPA.
Shisong Chen, Ziyu Zhou 0019, Zhixu Li, Yanghua Xiao, Xin Lin 0001, Xiaojun Meng, Jiansheng Wei, Kuien Liu
WSDM1
2026 Large Language Model Judged Self-Training for Named Entity Recognition
abstract
Self-training for Named Entity Recognition (NER) aims at identifying named entities and their types in the text using self-training to fully make use of the limited labeled data and a large amount of unlabeled data. The major challenge in self-training is confirmation bias where incorrect pseudo-labels increase errors. Many efforts have been made to address this challenge, but few labeled data limit their performance. In this paper, we introduce Large Language Model (LLM) into self-training to select high-quality pseudo-labels leveraging its rich knowledge and few-shot learning capability. Specifically, we design a comprehensive prompt to improve the judgment performance of LLM, where the prompt incorporates task rules mined by LLM itself to fully leverage labeled data. In addition, to reduce the impact of LLM's hallucinations, we adopt a collaborative pseudo-label selection based on combined confidence and calibration-guided probability smoothing. Our empirical study conducted on several NER datasets shows that our method outperforms state-of-the-art approaches. The code is available at https://github.com/cheniison/llm-judged-ST.
Shisong Chen, Jiaan Wang, Yanghua Xiao, Zhixu Li, Xin Lin 0001
WSDM1
2025 KUG: Joint Enhancement of Internal and External Knowledge for Retrieval-Augmented Generation
abstract
Query enhancement, a pivotal methodology in Retrieval-Augmented Generation (RAG) for addressing information scarcity in queries, has garnered increasing research attention. Nevertheless, existing approaches overlook the inherent distinctions between domain-specific knowledge and external factual sources during integration. To bridge this gap, we propose KUG (Knowledge-Update-Generation), a novel RAG framework that leverages internal knowledge semantics to ensure query enhancement efficacy, validates and dynamically updates knowledge representations using external evidence, and achieves systematic integration through knowledge graph embeddings. Extensive experiments on six standard BEIR benchmarks demonstrate that KUG outperforms the state-of-the-art methods, achieving an improvement of 1%-2% in recall metrics. Notably, the framework demonstrates significant performance gains in multi-hop reasoning tasks, advancing the development paradigm for RAG systems. The code will be public soon.
Shisong Chen, Shengkun Tu, Ziyi Du, Zhixu Li, Yanghua Xiao
CIKM2
2025 ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
abstract
Recent advances in large language models (LLMs) have demonstrated impressive capabilities in code-related tasks such as code generation and automated program repair. Despite their promising performance, most existing approaches for code repair suffer from high training costs or computationally expensive inference. Retrieval-augmented generation (RAG), with its efficient in-context learning paradigm, offers a more scalable alternative. However, conventional retrieval strategies, which are often based on holistic code-text embeddings, fail to capture the structural intricacies of code, resulting in suboptimal retrieval quality. To address the above limitations, we propose ReCode, a fine-grained retrieval-augmented in-context learning framework designed for accurate and efficient code repair. Specifically, ReCode introduces two key innovations: (1) an algorithm-aware retrieval strategy that narrows the search space using preliminary algorithm type predictions; and (2) a modular dual-encoder architecture that separately processes code and textual inputs, enabling fine-grained semantic matching between input and retrieved contexts. Furthermore, we propose RACodeBench, a new benchmark constructed from real-world user-submitted buggy code, which addresses the limitations of synthetic benchmarks and supports realistic evaluation. Experimental results on RACodeBench and competitive programming datasets demonstrate that ReCode achieves higher repair accuracy with significantly reduced inference cost, highlighting its practical value for real-world code repair scenarios.
Shisong Chen, Zhixu Li
CIKM2
2025 OVEL: Online Video Entity Linking
abstract
Recently, Multi-modal Entity Linking (MEL) has attracted increasing attention in the research community due to its significance in numerous multi-modal applications. Video, as a popular means of information transmission, has become prevalent in people’s daily lives. However, most existing MEL methods primarily focus on linking textual and visual mentions or offline videos’ mentions to entities in multi-modal knowledge bases, with limited efforts devoted to linking mentions within online video content. In this paper, we propose a task called Online Video Entity Linking (OVEL), aiming to establish connections between mentions in online videos and a knowledge base with high accuracy and timeliness. To facilitate the research works of (OVEL), we specifically concentrate on live delivery scenarios and construct a live delivery entity linking dataset called (LIVE). Besides, we propose an evaluation metric that considers robustness, timelessness, and accuracy. Furthermore, to effectively handle (OVEL) task, we leverage a memory block managed by a Large Language Model and retrieve entity candidates from the knowledge base to augment LLM performance on memory management. The experimental results prove the effectiveness and efficiency of our method.
Haiquan Zhao 0002, Xuwu Wang, Shisong Chen, Zhixu Li, Yanghua Xiao
COLING3
2024 A Hierarchy-aware Entity Alignment Method for Educational Knowledge Graphs
Anting Li, Shisong Chen, Zhixu Li, Jianfeng Qu, Zhiang Yue
DASFAA (4)2
2024 ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
abstract
Haiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong, Jiaan Wang, Kexin Huang, Tianle Gu, Yixu Wang, Jian Wang, Liang Dandan, Zhixu Li, Yan Teng, Yanghua Xiao, Yingchun Wang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiquan Zhao 0002, Shisong Chen, Shuqi Kong, Jiaan Wang, Tianle Gu, Yixu Wang, Dandan Liang, Zhixu Li, Yan Teng 0002, Yanghua Xiao, Yingchun Wang 0004
EMNLP3
2024 Generating Prompts in Latent Space for Rehearsal-free Continual Learning
abstract
Continual learning emerges as a framework that trains the model on a sequence of tasks without forgetting previously learned knowledge, which has been applied in multiple multimodal scenarios. Recently, prompt-based continual learning has achieved excellent domain adaptability and knowledge transfer through prompt generation. However, existing methods mainly focus on designing the architecture of a generator, neglecting the importance of providing effective guidance for training the generator. To address this issue, we propose Generating Prompts in Latent Space (GPLS), which considers prompts as latent variables to account for the uncertainty of prompt generation and aligns with the fact that prompts are inserted into the hidden layer outputs and exert an implicit influence on classification. GPLS adopts a trainable encoder to encode task and feature information into prompts with reparameterization technique, and provides refined and targeted guidance for the training process through the evidence lower bound (ELBO) related to Mahalanobis distance. Extensive experiments demonstrate that GPLS achieves state-of-the-art performance on various benchmarks. Our code is available at https://github.com/Hifipsysta/GPLS.
Shisong Chen, Jiayin Qi, Aimin Zhou
ACM Multimedia3
2021 Tackling Zero Pronoun Resolution and Non-Zero Coreference Resolution Jointly
abstract
Zero pronoun resolution aims at recognizing dropped pronouns and pointing out their anaphoric mentions, while non-zero coreference resolution targets at clustering mentions referring to the same entity.Existing efforts often deal with the two problems separately regardless of their close essential correlations.In this paper, we investigate the possibility of jointly solving zero pronoun resolution and coreference resolution via a novel end-to-end neural model.Specifically, we design a gapmasked self-attention model that encodes gaps and tokens in the same space, where gaps could capture valuable contextual information according to their surrounding tokens while tokens could maintain original sequential information without disturbance.Additionally, we also propose a two-stage interaction mechanism to make full use of the exclusive relationship between zero pronouns and mentions.Our empirical study conducted on the OntoNotes 5.0 Chinese dataset shows that our model could outperform corresponding stateof-the-art approaches on both tasks.
Shisong Chen, Binbin Gu, Jianfeng Qu, Zhixu Li, An Liu 0002, Lei Zhao 0001, Zhigang Chen 0003
CoNLL1