VLDB 2026 Research / reviewers in the wild / expert
Shengping Liu
dblp:21/5679
· DBLP profile ↗
29ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language ModelsabstractOmni-modal large language models (OLLMs) offer a promising end-to-end solution for slideenhanced speech recognition due to their inherent multimodal capabilities.However, we found a fundamental issue faced by OLLMs: Visual Interference, where models show a bias towards visible text over auditory signals, causing them to hallucinate slide content that was never spoken.To address this, we propose Visually-Anchored Policy Optimization (VAPO), which aims to reshape models' inference process to follow the human-like "Look-then-Listen" inference chain.Specifically, we design a temporally decoupled policy: the model first extracts visual priors in a block to serve as semantic anchors, then generates the transcription in an block.The policy is optimized via multi-objective reinforcement learning.Furthermore, we introduce SlideASR-Bench, a comprehensive benchmark designed to address the scarcity of entity-rich data, comprising a large-scale synthetic corpus for training and a challenging real-world test set for evaluation.We conduct extensive evaluations demonstrating that VAPO effectively eliminates visual interference and achieves state-of-the-art performance on SlideASR-Bench and public datasets, significantly reducing entity recognition errors in specialized domains. Rui Hu 0011, Delai Qiu, Shengping Liu, Jitao Sang 0001 |
ACL (1) | 4 |
| 2026 | FocalOrder: Focal Preference Optimization for Reading Order DetectionabstractFuyuan Liu, Dianyu Yu, He Ren, Nayu Liu, Xiaomian Kang, Delai Qiu, Fa Zhang, Genpeng Zhen, Shengping Liu, Liang Jiaen, Weihuang, Yining Wang, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fuyuan Liu, Dianyu Yu, Nayu Liu, Xiaomian Kang, Delai Qiu, Genpeng Zhen, Shengping Liu, Jiaen Liang, Junnan Zhu |
ACL (1) | 9 |
| 2026 | Comfort temperature assessment for honeybee colonies based on long-term monitoring
Yuntao Lu, Shijuan Li, Shengping Liu |
Expert Syst. Appl. | 6 |
| 2025 | Cracking Factual Knowledge: A Comprehensive Analysis of Degenerate Knowledge Neurons in Large Language ModelsabstractKnowledge neuron theory provides a key approach to understanding the mechanisms of factual knowledge in Large Language Models (LLMs), which suggests that facts are stored within multi-layer perceptron neurons. This paper further explores Degenerate Knowledge Neurons (DKNs), where distinct sets of neurons can store identical facts, but unlike simple redundancy, they also participate in storing other different facts. Despite the novelty and unique properties of this concept, it has not been rigorously defined and systematically studied. Our contributions are: (1) We pioneer the study of structures in knowledge neurons by analyzing weight connection patterns, providing a comprehensive definition of DKNs from both functional and structural aspects. (2) Based on this definition, we develop the Neuronal Topology Clustering method, leading to a more accurate DKN identification. (3) We demonstrate the practical applications of DKNs in two aspects: guiding LLMs to learn new knowledge and relating to LLMs’ robustness against input errors. Yubo Chen 0001, Shengping Liu, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 5 |
| 2025 | Awakening Augmented Generation: Learning to Awaken Internal Knowledge of Large Language Models for Question AnsweringabstractRetrieval-Augmented-Generation and Generation-Augmented-Generation have been proposed to enhance the knowledge required for question answering with Large Language Models (LLMs) by leveraging richer context. However, the former relies on external resources, and both require incorporating explicit documents into the context, which increases execution costs and susceptibility to noise data during inference. Recent works indicate that LLMs model rich knowledge, but it is often not effectively activated and awakened. Inspired by this, we propose a novel knowledge-augmented framework, Awakening-Augmented-Generation (AAG), which mimics the human ability to answer questions using only thinking and recalling to compensate for knowledge gaps, thereby awaking relevant knowledge in LLMs without relying on external resources. AAG consists of two key components for awakening richer context. Explicit awakening fine-tunes a context generator to create a synthetic, compressed document that functions as symbolic context. Implicit awakening utilizes a hypernetwork to generate adapters based on the question and synthetic document, which are inserted into LLMs to serve as parameter context. Experimental results on three datasets demonstrate that AAG exhibits significant advantages in both open-domain and closed-book settings, as well as in out-of-distribution generalization. Our code will be available at https://github.com/Xnhyacinth/IAG. Huanxuan Liao, Shizhu He, Yuanzhe Zhang, Shengping Liu, Kang Liu 0001, Jun Zhao 0001 |
COLING | 5 |
| 2025 | Prompt robust large language model for Chinese medical named entity recognition
Yubo Chen 0001, Baoli Zhang, Zhuoran Jin, Zhengyuan Cai, Yingzheng Wang, Delai Qiu, Shengping Liu, Jun Zhao 0001 |
Inf. Process. Manag. | 8 |
| 2024 | Oasis: Data Curation and Assessment System for Pretraining of Large Language Models
Tong Zhou 0014, Yubo Chen 0001, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
IJCAI | 5 |
| 2024 | Optoelectronic Computing Evaluation and Deployment Platform Based on a 256-MAC Silicon Photonic ChipabstractThe deceleration of Moore's Law has led to increasing difficulties in advancing the computational speed and power efficiency of Complementary-Metal-Oxide-Semiconductor (CMOS) chips. As a solution to this challenge, optical computing emerges as a promising technology, boasting low energy consumption, high processing speed, and extensive bandwidth. Yet, a critical obstacle remains: the absence of a co-simulation platform that incorporates both photonic chips and peripheral electrical circuits. This paper addresses this gap by introducing a hybrid optoelectronic computing evaluation and deployment platform utilizing Simulink tools. Based on the measured data from the silicon optical computing chip, we have deployed an image filtering algorithm and a convolutional neural network onto this platform. The optical computing chip achieves an accuracy of 86.4% on the ImageNet image dataset. Through evaluation, we have identified the most substantial impacts on calculation results. To achieve an image classification accuracy of 80%, the signal-to-noise ratio (SNR) of the low-speed DAC must be a minimum of 52 dB. These findings provide crucial insights into the optimization of optical computing systems. Likai Li, Yichuan Bai, Shengping Liu, Sunan He, Yaqing Li, Yuan Du |
ISCAS | 3 |
| 2024 | From Instance Training to Instruction Learning: Task Adapters Generation from InstructionsabstractLarge language models (LLMs) have acquired the ability to solve general tasks by utilizing instruction finetuning (IFT). However, IFT still relies heavily on instance training of extensive task data, which greatly limits the adaptability of LLMs to real-world scenarios where labeled task instances are scarce and broader task generalization becomes paramount. Contrary to LLMs, humans acquire skills and complete tasks not merely through repeated practice but also by understanding and following instructional guidelines. This paper is dedicated to simulating human learning to address the shortcomings of instance training, focusing on instruction learning to enhance cross-task generalization. Within this context, we introduce Task Adapters Generation from Instructions (TAGI), which automatically constructs the task-specific model in a parameter generation manner based on the given task instructions without retraining for unseen tasks. Specifically, we utilize knowledge distillation to enhance the consistency between TAGI developed through Learning with Instruction and task-specific models developed through Training with Instance, by aligning the labels, output logits, and adapter parameters between them. TAGI is endowed with cross-task generalization capabilities through a two-stage training process that includes hypernetwork pretraining and finetuning. We evaluate TAGI on the Super-Natural Instructions and P3 datasets. The experimental results demonstrate that TAGI can match or even outperform traditional meta-trained models and other hypernetwork models, while significantly reducing computational requirements. Our code will be available at https://github.com/Xnhyacinth/TAGI. Huanxuan Liao, Shizhu He, Yuanzhe Zhang, Yanchao Hao, Shengping Liu, Kang Liu 0001, Jun Zhao 0001 |
NeurIPS | 6 |
| 2024 | Large Language Models With Holistically Thought Could Be Better Doctors
Yixuan Weng, Bin Li 0083, Minjun Zhu, Bin Sun 0001, Shizhu He, Shengping Liu, Kang Liu 0001, Shutao Li 0001, Jun Zhao 0001 |
NLPCC (2) | 7 |
| 2021 | Automatic ICD Coding via Interactive Shared Representation Networks with Self-distillation MechanismabstractTong Zhou, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao, Kun Niu, Weifeng Chong, Shengping Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tong Zhou 0014, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Kun Niu, Weifeng Chong, Shengping Liu |
ACL/IJCNLP (1) | 8 |
| 2021 | Biomedical Concept Normalization by Leveraging HypernymsabstractBiomedical Concept Normalization (BCN) is widely used in biomedical text processing as a fundamental module.Owing to numerous surface variants of biomedical concepts, BCN still remains challenging and unsolved.In this paper, we exploit biomedical concept hypernyms to facilitate BCN.We propose Biomedical Concept Normalizer with Hypernyms (BCNH), a novel framework that adopts list-wise training to make use of both hypernyms and synonyms, and also employs norm constraint on the representation of hypernym-hyponym entity pairs.The experimental results show that BCNH outperform the previous state-of-the-art model on the NCBI dataset. Yuanzhe Zhang, Kang Liu 0001, Jun Zhao 0001, Shengping Liu |
EMNLP (1) | 6 |
| 2021 | Incorporate Lexicon into Self-training: A Distantly Supervised Chinese Medical NER
Zhen Gan, Zhucong Li, Baoli Zhang, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Shengping Liu |
NLPCC (1) | 9 |
| 2021 | Clinical decision-making framework against over-testing based on modeling implicit evaluation criteria
Yang Yang 0041, Hongxing Huo, Jingchi Jiang, Xuemei Sun, Yi Guan, Xitong Guo, Shengping Liu |
J. Biomed. Informatics | 8 |
| 2021 | A Unified Shared-Private Network with Denoising for Dialogue State Tracking
Qingbin Liu, Shizhu He, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
J. Comput. Sci. Technol. | 4 |
| 2020 | HyperCore: Hyperbolic and Co-graph Representation for Automatic ICD CodingabstractThe International Classification of Diseases (ICD) provides a standardized way for classifying diseases, which endows each disease with a unique code.ICD coding aims to assign proper ICD codes to a medical record.Since manual coding is very laborious and prone to errors, many methods have been proposed for the automatic ICD coding task.However, most of existing methods independently predict each code, ignoring two important characteristics: Code Hierarchy and Code Co-occurrence.In this paper, we propose a Hyperbolic and Co-graph Representation method (HyperCore) to address the above problem.Specifically, we propose a hyperbolic representation method to leverage the code hierarchy.Moreover, we propose a graph convolutional network to utilize the code co-occurrence.Experimental results on two widely used datasets demonstrate that our proposed model outperforms previous state-ofthe-art methods. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Shengping Liu, Weifeng Chong |
ACL | 5 |
| 2020 | MIE: A Medical Information Extractor towards Medical DialoguesabstractElectronic Medical Records (EMRs) have become key components of modern medical care systems. Despite the merits of EMRs, many doctors suffer from writing them, which is time-consuming and tedious. We believe that automatically converting medical dialogues to EMRs can greatly reduce the burdens of doctors, and extracting information from medical dialogues is an essential step. To this end, we annotate online medical consultation dialogues in a window-sliding style, which is much easier than the sequential labeling annotation. We then propose a Medical Information Extractor (MIE) towards medical dialogues. MIE is able to extract mentioned symptoms, surgeries, tests, other information and their corresponding status. To tackle the particular challenges of the task, MIE uses a deep matching architecture, taking dialogue turn-interaction into account. The experimental results demonstrate MIE is a promising solution to extract medical information from doctor-patient dialogues. Yuanzhe Zhang, Zhongtao Jiang, Tao Zhang 0097, Shiwan Liu, Jiarun Cao, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
ACL | 7 |
| 2019 | CBOWRA: A Representation Learning Approach for Medication Anomaly DetectionabstractElectronic health record is an important source for clinical researches and applications, and errors inevitably occur in the data, which lead to severe damages to both patients and hospital services. One of such errors is the mismatch between diagnose and prescription, which we address as “medication anomaly” in the paper, and clinicians used to manually identify and correct them. With the development of machine learning techniques, researchers are able to train specific model for the task, but the process still requires expert knowledge to construct proper features, and few semantic relations are considered. In this paper, we propose a simple, yet effective detection method that tackles the problem by detecting the semantic inconsistency between diagnoses and prescriptions. Unlike traditional outlier or anomaly detection, the scheme uses continuous bag of words to construct the semantic connection between specific central words and their surrounding context. The detection of medication anomaly is transformed into identifying the least possible central word based on given context. To help distinguish the anomaly from normal context, we also incorporate a ranking accumulation strategy. The experiments were conducted on two real hospital electronic medical records, and the topN accuracy of the proposed method increased by 3.91 to 10.91% and 0.68 to 2.13% on the datasets, respectively, which is highly competitive to other traditional machine learning-based approaches. Zhiyuan Ma 0001, Yangming Zhou, Shengping Liu, Ju Gao, Wen Du |
BIBM | 5 |
| 2019 | Leverage Lexical Knowledge for Chinese Named Entity Recognition via Collaborative Graph NetworkabstractDianbo Sui, Yubo Chen, Kang Liu, Jun Zhao, Shengping Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dianbo Sui, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Shengping Liu |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Learning the Extraction Order of Multiple Relational Facts in a Sentence with Reinforcement LearningabstractXiangrong Zeng, Shizhu He, Daojian Zeng, Kang Liu, Shengping Liu, Jun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiangrong Zeng, Shizhu He, Daojian Zeng, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Adversarial Transfer Learning for Chinese Named Entity Recognition with Self-Attention MechanismabstractNamed entity recognition (NER) is an important task in natural language processing area, which needs to determine entities boundaries and classify them into pre-defined categories.For Chinese NER task, there is only a very small amount of annotated data available.Chinese NER task and Chinese word segmentation (CWS) task have many similar word boundaries.There are also specificities in each task.However, existing methods for Chinese NER either do not exploit word boundary information from CWS or cannot filter the specific information of CWS.In this paper, we propose a novel adversarial transfer learning framework to make full use of task-shared boundaries information and prevent the taskspecific features of CWS.Besides, since arbitrary character can provide important cues when predicting entity type, we exploit selfattention to explicitly capture long range dependencies between two tokens.Experimental results on two different widely used datasets show that our proposed model significantly and consistently outperforms other state-ofthe-art methods. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Shengping Liu |
EMNLP | 5 |
| 2009 | iSMART: Ontology-based Semantic Query of CDA Documents
Shengping Liu, Yuan Ni, Jing Mei, Guo Tong Xie, Gang Hu 0001, Haifeng Liu 0005, Xueqiao Hou |
AMIA | 1 |
| 2009 | A Practical Approach for Scalable Conjunctive Query Answering on Acyclic EL+\mathcal{EL}^+ Knowledge Base
Jing Mei, Shengping Liu, Guo Tong Xie, Aditya Kalyanpur, Achille Fokoue, Yuan Ni |
ISWC | 2 |
| 2008 | Supporting Ontology-Based Dynamic Property and Classification in WebSphere Metadata Server
Shengping Liu, Yang Yang 0041, Guo Tong Xie, Chen Wang 0020, Cassio Dos Santos, Robert J. Schloss, Kevin Shank, John Colgrave |
ISWC | 1 |
| 2006 | Towards a Complete OWL Ontology Benchmark
Li Ma 0002, Yang Yang 0041, Zhaoming Qiu, Guo Tong Xie, Shengping Liu |
ESWC | 6 |
| 2005 | Representation and Reasoning on RBAC: A Description Logic Approach
Chen Zhao 0001, NuerMaimaiti Heilili, Shengping Liu, Zuoquan Lin |
ICTAC | 3 |
| 2004 | A Real-Time Information Gathering Agent Based on Ontology
Shengping Liu, Zuoquan Lin, Cen Wu |
WAIM | 2 |
| 2003 | An Evaluation on Feature Selection for Text Clustering
Shengping Liu, Zheng Chen 0001, Wei-Ying Ma |
ICML | 2 |
| 2003 | Building a web thesaurus from web link structureabstractThesaurus has been widely used in many applications, including information retrieval, natural language processing, and question answering. In this paper, we propose a novel approach to automatically constructing a domain-specific thesaurus from the Web using link structure information. The proposed approach is able to identify new terms and reflect the latest relationship between terms as the Web evolves. First, a set of high quality and representative websites of a specific domain is selected. After filtering out navigational links, link analysis is applied to each website to obtain its content structure. Finally, the thesaurus is constructed by merging the content structures of the selected websites. The experimental results on automatic query expansion based on our constructed thesaurus show 20% improvement in search precision compared to the baseline. Zheng Chen 0001, Shengping Liu, Wenyin Liu, Geguang Pu, Wei-Ying Ma |
SIGIR | 2 |