VLDB 2026 Research / reviewers in the wild / expert
Shuaimin Li
dblp:228/4334
· DBLP profile ↗
16ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-8368-916XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author DebatesabstractExisting paper review methods often rely on superficial manuscript features or directly on large language models (LLMs), which are prone to hallucinations, biased scoring, and limited reasoning capabilities. Moreover, these methods often fail to capture the complex argumentative reasoning and negotiation dynamics inherent in reviewer-author interactions. To address these limitations, we propose ReViewGraph (Reviewer-Author Debates Graph Reasoner), a novel framework that performs heterogeneous graph reasoning over LLM-simulated multi-round reviewer-author debates. In our approach, reviewer-author exchanges are simulated through LLM-based multi-agent collaboration. Diverse opinion relations (e.g., acceptance, rejection, clarification, and compromise) are then explicitly extracted and encoded as typed edges within a heterogeneous interaction graph. By applying graph neural networks to reason over these structured debate graphs, ReViewGraph captures fine-grained argumentative dynamics and enables more informed review decisions. Extensive experiments on three datasets demonstrate that ReViewGraph outperforms strong baselines with an average relative improvement of 15.73%, underscoring the value of modeling detailed reviewer–author debate structures. Shuaimin Li, Liyang Fan, Yufang Lin, Xian Wei, Shiwen Ni, Hamid Alinejad-Rokny, Min Yang 0007 |
AAAI | 1 |
| 2026 | Self-SoftCoT: A Self-Consistent Framework via Position-Aware Latent Space Reinforcement LearningabstractWhile Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) to tackle complex tasks, its reliance on discrete token decoding imposes an inherent Discreteness Bottleneck, limiting expressiveness within a restricted vocabulary space.Existing continuous reasoning approaches, such as SoftCoT (Xu et al., 2025), mitigate this but typically rely on external auxiliary models, resulting in complex deployment and fractured inference pipelines.To address these challenges, we propose Self-SoftCoT, a self-contained framework that enables a frozen LLM to internally generate and consume latent thoughts without external assistants.By establishing a singlestream "Thinking → Speaking" closed-loop, we decouple latent planning from explicit generation.Furthermore, we adopt Group Sequence Policy Optimization (GSPO) to stabilize learning and employ Position-Aware Independent Projection to mitigate representation homogenization.Experimental results on five reasoning benchmarks demonstrate that our method significantly improves the reasoning performance of frozen LLMs.Specifically, our Qwen2.5-basedmodel (Yang et al., 2024) uses only N = 2 soft tokens to outperform the Soft-CoT baseline (N = 4), improving the average accuracy from 75.06% to 78.42%.Similarly, LLaMA-3.1 (Llama Team, 2024) performance increases from 70.52% to 74.55% 1 . Liangliang Dong, Lianlei Shan, Shuaimin Li |
ACL (1) | 3 |
| 2026 | WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationabstractKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang, Jiasheng Si, Bing Xu, Kehai Chen, Muyun Yang, Wenpeng Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kunpeng Kang, Shuaimin Li, Jiasheng Si, Kehai Chen, Muyun Yang, Wenpeng Lu |
ACL (1) | 2 |
| 2026 | VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization ModelsabstractText-to-visualization (text-to-vis) models for tabular data have become essential tools in the era of big data, enabling users to generate visualizations and make data-driven decisions through natural language queries (NLQs). Despite their growing adoption, the security vulnerabilities of these models remain largely unexplored. To address this gap, we propose VisPoison, a backdoor attack framework that realistically simulates three types of attacks on text-to-vis models via data poisoning: data exposure, misleading visualizations, and denial-of-service (DoS). Specifically, VisPoison introduces two types of stealthy triggers to enable both proactive and passive backdoor activations. Proactive triggers are deliberately inserted by attackers using rare-word patterns to extract sensitive information, whereas passive triggers are unintentionally activated by users through first-word prompts, resulting in visualization errors or DoS failures. To support these triggers, we craft specialized payloads for visualization queries that allow compromised models to function normally on benign inputs while producing malicious outputs in the presence of triggers. Extensive evaluations on both trainable and in-context learning (ICL)-based text-to-vis models show that VisPoison achieves attack success rates exceeding 90\%, exposing serious vulnerabilities. Additionally, existing defense strategies reveal limited effectiveness against VisPoison, underscoring the urgent need for more robust and security-aware text-to-vis systems to safeguard human-data interaction. Shuaimin Li, Chen Zhang 0013, Xuanang Chen, Anni Peng, Zhuoyue Wan, Yuanfeng Song, Shiwen Ni, Min Yang 0007, Raymond Chi-Wing Wong |
ICDE | 1 |
| 2026 | OsmT: Bridging Openstreetmap Queries and Natural Language With Open-Source Tag-Aware Language ModelsabstractBridging natural language and structured query languages is a long-standing challenge in the database community. While recent advances in language models have shown promise in this direction, existing solutions often rely on large-scale closed-source models that suffer from high inference costs, limited transparency, and lack of adaptability for lightweight deployment. In this paper, we present OsmT, an open-source tag-aware language model specifically designed to bridge natural language and Overpass Query Language (OverpassQL), a structured query language for accessing large-scale OpenStreetMap (OSM) data. To enhance the accuracy and structural validity of generated queries, we introduce a Tag Retrieval Augmentation (TRA) mechanism that incorporates contextually relevant tag knowledge into the generation process. This mechanism is designed to capture the hierarchical and relational dependencies present in the OSM database, addressing the topological complexity inherent in geospatial query formulation. In addition, we define a reverse task, OverpassQL-to-Text, which translates structured queries into natural language explanations to support query interpretation and improve user accessibility. We evaluate OsmT on a public benchmark against strong baselines and observe consistent improvements in both query generation and interpretation. Despite using significantly fewer parameters, our model achieves competitive accuracy, demonstrating the effectiveness of open-source pre-trained language models in bridging natural language and structured query languages within schema-rich geospatial environments. Zhuoyue Wan, Chen Zhang 0013, Yuanfeng Song, Shuaimin Li, Ruiqiang Xiao, Xiaoyong Wei, Raymond Chi-Wing Wong |
ICDE | 5 |
| 2025 | RoDEval: A Robust Word Sense Disambiguation Evaluation Framework for Large Language ModelsabstractAccurately evaluating the word sense disambiguation (WSD) capabilities of large language models (LLMs) remains challenging, as existing studies primarily rely on single-task evaluations and classification-based metrics that overlook the fundamental differences between generative LLMs and traditional classification models.To bridge this gap, we propose RoDEval, the first comprehensive evaluation framework specifically tailored for assessing LLMbased WSD methods.RoDEval introduces four novel metrics: Disambiguation Scope, Disambiguation Robustness, Disambiguation Reliability, and Definition Generation Quality Score, enabling a multifaceted evaluation of LLMs' WSD capabilities.Experimental results using RoDEval across five mainstream LLMs uncover significant limitations in their WSD performance.Specifically, incorrect definition selections in multiple-choice WSD tasks stem not from simple neglect or forget of correct options, but rather from incomplete acquisition of the all senses for polysemous words.Instead, disambiguation reliability is often compromised by the models' persistent overconfidence.In addition, inherent biases continue to affect performance, and scaling up model parameters alone fails to meaningfully enhance their ability to generate accurate sense definitions.These findings provide actionable insights for enhancing LLMs' WSD capabilities.The source code and evaluation scripts are open-sourced at https: //github.com/DayDream405/RoDEval. Shuaimin Li, Yishuo Li, Kunpeng Kang, Wenpeng Lu |
EMNLP | 2 |
| 2025 | MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationabstractKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng, Shuaimin Li, Bing Xu, Muyun Yang, Xinxiao Qiao, Wenpeng Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qian Liu 0012, Chaoqun Zheng, Shuaimin Li, Muyun Yang, Xinxiao Qiao, Wenpeng Lu |
EMNLP | 5 |
| 2025 | DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data VisualizationabstractData visualization (DV) is the fundamental and premise tool to improve the efficiency in conveying the insights behind the big data, which has been widely accepted in existing data-driven world. Task automation in DV, such as converting natural language queries to visualizations (i.e., text-to-vis), gener-ating explanations from visualizations (i.e., vis-to-text), answering DV-related questions in free form (i.e. Fe VisQA), and explicating tabular data (i.e., table-to-text), is vital for advancing the field. Despite their potential, the application of pre-trained language models (PLMs) like T5 and BERT in DV has been limited by high costs and challenges in handling cross-modal information, leading to few studies on PLMs for DV. We introduce Data VisT5, a novel PLM tailored for DV that enhances the T5 architecture through a hybrid objective pre-training and multi-task fine-tuning strategy, integrating text and DV datasets to effectively interpret cross-modal semantics. Extensive evaluations on public datasets show that Data VisT5 consistently outperforms current state-of-the-art models and higher-parameter Large Language Models (LLMs) on various DV-related tasks. We anticipate that Data VisT5 will not only inspire further research on vertical PLMs but also expand the range of applications for PLMs. Zhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Zhang 0013, Raymond Chi-Wing Wong |
ICDE | 3 |
| 2025 | Zero-shot neural architecture search with weighted response correlation
Kun Jing, Luoyu Chen, Jungang Xu, Jianwei Tai, Shuaimin Li |
Neurocomputing | 6 |
| 2025 | prompt4vis: prompting large language models with example mining for tabular data visualizationabstractAbstract We are currently in the epoch of Large Language Models (LLMs), which have transformed numerous technological domains within the database community. In this paper, we examine the application of LLMs in text-to-visualization (text-to-vis). The advancement of natural language processing technologies has made natural language interfaces more accessible and intuitive for visualizing tabular data. However, despite utilizing advanced neural network architectures, current methods such as Seq2Vis, ncNet, and RGVisNet for transforming natural language queries into DV commands still underperform, indicating significant room for improvement. In this paper, we introduce Prompt4Vis , a novel framework that leverages LLMs and In-context learning to enhance the generation of data visualizations from natural language. Given that In-context learning’s effectiveness is highly dependent on the selection of examples, it is critical to optimize this aspect. Additionally, encoding the full database schema of a query is not only costly but can also lead to inaccuracies. This framework includes two main components: (1) an example mining module that identifies highly effective examples to enhance In-context learning capabilities for text-to-vis applications, and (2) a schema filtering module designed to streamline database schemas. Comprehensive testing on the NVBench dataset has shown that Prompt4Vis significantly outperforms the current state-of-the-art model, RGVisNet, by approximately 35.9% on development sets and 71.3% on test sets. To the best of our knowledge, Prompt4Vis is the first framework to incorporate In-context learning for enhancing text-to-vis, marking a pioneering step in the domain. Shuaimin Li, Xuanang Chen, Yuanfeng Song, Yunze Song, Chen Zhang 0013, Lei Chen 0002 |
VLDB J. | 1 |
| 2024 | Correction to: HierMDS: a hierarchical multi-document summarization model with global-local document dependencies
Shuaimin Li, Jungang Xu |
Neural Comput. Appl. | 1 |
| 2023 | A Novel Clinical Trial Prediction-Based Factual Inconsistency Detection Approach for Medical Text SummarizationabstractMost existing works of factual inconsistency detection focus on text summarization of generic articles. In this paper, a clinical trial prediction-based factual inconsistency detection approach is proposed for medical text summarization. Inspired by the fact that medical articles related to a clinical trial can give some evidence to predict the outcome of this trial, we believe that a factual consistent summary of the medical articles can also accurately predict the outcome of the corresponding clinical trial. Therefore, we first gather a novel Clinical Trial Prediction-based summarization (CTPSum) dataset, which is a collection of the outcomes of the clinical trials together with the related medical articles and summaries. We then propose a keyword-aware classification model to predict the outcome (successful/failed) of a clinical trial. If the predicted outcome is correct, the generated summary is considered to be factually consistent with the source document, otherwise, it is labeled as factually inconsistent. To evaluate the proposed approach, we further collect a Fact Inconsistency Detection dataset in the Medical Domain (FIDMD), which includes summaries, medical articles, and binary labels indicating factual consistency or inconsistency. Experimental results demonstrate the effectiveness of the proposed approach, specifically, the clinical trial prediction-based factual inconsistency detection approach outperforms several NLI-based factual inconsistency detection methods on the FIDMD dataset. Shuaimin Li, Jungang Xu |
IJCNN | 1 |
| 2023 | MRC-Sum: An MRC framework for extractive summarization of academic articles in natural sciences and medicine
Shuaimin Li, Jungang Xu |
Inf. Process. Manag. | 1 |
| 2022 | KAAS: A Keyword-Aware Attention Abstractive Summarization Model for Scientific Articles
Shuaimin Li, Jungang Xu |
DASFAA (3) | 1 |
| 2021 | TRANSREGEX: Multi-modal Regular Expression Synthesis by Generate-and-RepairabstractSince regular expressions (abbrev. regexes) are difficult to understand and compose, automatically generating regexes has been an important research problem. This paper introduces TransRegex, for automatically constructing regexes from both natural language descriptions and examples. To the best of our knowledge, TransRegex is the first to treat the NLP-and-example-based regex synthesis problem as the problem of NLP-based synthesis with regex repair. For this purpose, we present novel algorithms for both NLP-based synthesis and regex repair. We evaluate TransRegex with ten relevant state-of-the-art tools on three publicly available datasets. The evaluation results demonstrate that the accuracy of our TransRegex is 17.4%, 35.8% and 38.9% higher than that of NLP-based approaches on the three datasets, respectively. Furthermore, TransRegex can achieve higher accuracy than the state-of-the-art multi-modal techniques with 10% to 30% higher accuracy on all three datasets. The evaluation results also indicate TransRegex utilizing natural language and examples in a more effective way. Yeting Li, Shuaimin Li, Zhiwu Xu 0001, Jialun Cao, Haiming Chen 0001, Shing-Chi Cheung |
ICSE | 2 |
| 2021 | A two-step abstractive summarization model with asynchronous and enriched-information decoding
Shuaimin Li, Jungang Xu |
Neural Comput. Appl. | 1 |