VLDB 2026 Research / reviewers in the wild / expert
Shengyuan Bai
dblp:372/2052
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-9922-0278ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 47% Knowledge representation and reasoning · 26% Question answering and dialogue systems · 15% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
1.0 | 1 | 2026 | Counterfactual-based Cognitive Alignment In-Context Learning for Relation Extraction · AAAI 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning |
1.0 | 1 | 2026 | Counterfactual-based Cognitive Alignment In-Context Learning for Relation Extraction · AAAI 2026 |
Natural language and speech › Language models and text generation
in-context learning |
1.0 | 1 | 2026 | Counterfactual-based Cognitive Alignment In-Context Learning for Relation Extraction · AAAI 2026 |
Natural language and speech › Information extraction and text analysis
relation extraction |
1.0 | 1 | 2026 | Counterfactual-based Cognitive Alignment In-Context Learning for Relation Extraction · AAAI 2026 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Enhancing NLU in Large Language Models Using Adversarial Noisy Instruction Tuning · AAAI 2025 |
Natural language and speech › Language models and text generation
natural language understanding |
0.9 | 1 | 2025 | Enhancing NLU in Large Language Models Using Adversarial Noisy Instruction Tuning · AAAI 2025 |
Natural language and speech › Question answering and dialogue systems › open-domain dialogue
role-playing dialogue agents |
0.9 | 1 | 2025 | Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agents · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
0.9 | 1 | 2025 | Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agents · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
example selection · 1.0counterfactual generation · 1.0cognitive alignment · 1.0supervised fine-tuning · 0.9semantic distortion quantification · 0.9prompt construction · 0.9low-resource data construction · 0.9adversarial training · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Counterfactual-based Cognitive Alignment In-Context Learning for Relation ExtractionabstractLarge Language Models (LLMs) have demonstrated remarkable In-Context learning (ICL) capabilities for relation extraction (RE). While ICL has shown promise in RE tasks, current approaches face challenges in example selection and utilization. These challenges stem from the misalignment between example selection methods and LLMs' inherent cognitive processing mechanisms, particularly in pattern recognition and relational reasoning. To address these limitations, we propose Counterfactual Cognitive Alignment (CCA), a novel framework that systematically enhances ICL performance in RE by aligning example selection with cognitive principles underlying human relational reasoning. The framework incorporates a cognitive-inspired counterfactual generation mechanism that creates semantically diverse yet relationally coherent examples, mirroring human "what-if" reasoning processes. Additionally, it employs a cognitive alignment approach that integrates structural identification features with semantic understanding to better align with LLMs cognitive processing patterns. Extensive experiments across multiple RE benchmarks reveal the effectiveness of our cognitive alignment approach through the synergistic integration of counterfactual reasoning and cognitively-guided selection. Qibin Li, Shengyuan Bai, Nai Zhou, Nianmin Yao |
AAAI | 2 |
| 2026 | PriV2I: Privacy-preserving V2I authentication protocol with fine-grained access controlabstractAs vehicular ad hoc networks (VANETs) increase in size and complexity, ensuring secure, flexible, and privacy-preserving vehicle-to-infrastructure (V2I) authentication remains a major challenge. Existing protocols often focus solely on identity verification, overlooking the need for access control based on vehicle attributes. Furthermore, vehicles must obtain authentication credentials from various trusted entities, including automakers, regulators, and government agencies. However, the absence of a unified credential issuance mechanism introduces fragmentation and inconsistencies during the registration process. To address these issues, we propose a V2I authentication protocol, called PriV2I, that integrates distributed credential issuance, attribute-based access control, and strong anonymity guarantees. During vehicle registration, our approach uses Shamir’s Secret Sharing with a threshold t of n across multiple certification authorities (CAs) to consolidate credentials. A vehicle credential can only be issued by a predefined threshold number of CAs, enhancing security and flexibility. Within the authentication protocol, Pointcheval-Sanders (PS) signatures enable fine-grained access control based on vehicle attributes such as type and role. Meanwhile, noninteractive zero-knowledge proofs protect identity privacy by allowing vehicles to prove credential possession and policy compliance without revealing sensitive information. The proposed scheme also supports batch authentication at Roadside Units (RSUs) to efficiently handle high-density environments and includes a comprehensive revocation mechanism to trace and revoke malicious vehicles promptly and securely. In our implementation, the computation cost during the authentication phase is 75.58 ms. The communication overhead per authentication exchange is 992 bytes across two messages. Overall, the protocol provides a secure, scalable, and privacy-preserving solution tailored to modern VANET environments. Zhengze Liu, Nianmin Yao, Shengyuan Bai, Tengyi Mai |
Ad Hoc Networks | 3 |
| 2026 | Causally graph-guided counterfactual analysis to biomedical named entity recognition
Qibin Li, Shengyuan Bai, Nai Zhou, Nianmin Yao |
Expert Syst. Appl. | 2 |
| 2025 | Enhancing NLU in Large Language Models Using Adversarial Noisy Instruction TuningabstractInstruction tuning has emerged as an effective approach that notably improves large language models (LLMs) performance, showing particular promise in natural language generation tasks by producing more diverse, coherent, and task-relevant outputs. However, extending instruction tuning to natural language understanding (NLU) tasks presents significant challenges, primarily due to the difficulty in achieving high-precision responses and the scarcity of large-scale, high-quality instruction data necessary for effective tuning. In this work, we introduce Adversarial Noisy Instruction Tuning (ANIT) to improve NLU performance on LLMs. First, we leverage low-resource techniques to construct noisy instruction datasets. Second, we employ semantic distortion-aware techniques to quantify the intensity of noise within these instructions. Last, we devise an adversarial training method that incorporates a noise response strategy to achieve noisy instruction tuning. ANIT enhances LLMs capability to detect and accommodate semantic distortions in noisy instructions, thereby augmenting their comprehension of task objectives and ability to generate more accurate responses. We evaluate our approach across diverse noisy instructions and semantic distortion quantification methods on multiple NLU tasks. Comprehensive empirical results demonstrate that our method consistently outperforms existing approaches across various experimental settings. Shengyuan Bai, Qibin Li, Nai Zhou, Nianmin Yao |
AAAI | 1 |
| 2025 | Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational AgentsabstractLarge Language Models (LLMs) have demonstrated significant advancements in various fields, notably in Role-Playing Conversational Agents (RPCAs).However, when confronted with role-specific professional inquiries, LLMsbased RPCAs tend to underperform due to their excessive emphasis on the conversational abilities of characters rather than effectively invoking and integrating relevant expert knowledge.This often results in inaccurate responses.We refer to this phenomenon as the "Knowledge Misalignment" which underscores the limitations of RPCAs in integrating expert knowledge.To mitigate this issue, we have introduced an Anchoring-Guidance Fine-Tuning (AnGFT) Framework into the RPCAs' training process.This involves initially linking the Anchoring-Based System Prompt (ASP) with the LLM's relevant expert domains through diverse prompt construction strategies and supervised fine-tuning (SFT).Following the roleplay enriched SFT, the integration of ASP enables LLMs to better associate with relevant expert knowledge, thus enhancing their response capabilities in role-specific expert domains.Moreover, we have developed four comprehensive metrics-helpfulness, thoroughness, credibility, and feasibility-to evaluate the proficiency of RPCAs in responding to professional questions.Our method was tested across four professional fields, and the experimental outcomes suggest that the proposed AnGFT Framework substantially improves the RPCAs' performance in handling role-specific professional queries, while preserving their robust role-playing abilities. Qibin Li, Shengyuan Bai, Nianmin Yao, Kaili Sun, Baoxun Wang |
EMNLP | 3 |
| 2025 | Enhancing entity and relation extraction with dynamic hard negative augmentation framework
Qibin Li, Shengyuan Bai, Nai Zhou, Nianmin Yao |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | ImmuFold: High-Accuracy Antibody Structure Prediction with Efficient NetworkabstractAntibody structure prediction is a critical task in immunological research and therapeutic antibody development. Despite advances in prediction methods, contemporary approaches still face formidable challenges, particularly in accurately modeling Complementarity-determining regions (CDRs). Furthermore, current prediction time costs, typically on the order of minutes, preclude large-scale structure prediction and screening. In this work, we present ImmuFold, a novel deep-learning approach that achieves second-level performance in antibody structure prediction. ImmuFold integrates ImmuBERT, a 650M antibody language model pre-trained on hundreds of millions of natural antibody sequences, with a structure prediction network that directly predicts all-atom structure, encompassing both main chain and side chains. ImmuFold outperforms current methods, including IgFold and AlphaFold2, generating higher-quality antibody structures in approximately one second. Comparative analysis of the antibody binding task demonstrates the superior representational capabilities of ImmuBERT relative to existing language models, a crucial factor underpinning the efficacy of ImmuFold. Shengyuan Bai, Zijing Liu, Jiying Zhang, Yu Li 0003 |
BIBM | 1 |
| 2024 | Enhancing Biomedical NER with Adversarial Selective TrainingabstractLarge language models (LLMs) have significantly impacted the field of natural language processing (NLP). However, due to the limited domain specificity of the training data and the model’s constrained ability to generalize across complex biomedical data, LLMs continue to encounter challenges related to prediction bias and low generalization in biomedical named entity recognition (BioNER). In this work, we set out to improve the recognition and generalization capabilities of LLMs in BioNER through an Adversarial Selective Training (AST) method. Our method maximizes the adversarial loss to obtain the importance ranking of weights, which guides the model to selectively train to generate counterfactual examples. This strategy aims to force the model to explore the amount of information in the latent space to extract entities, thereby improving the performance of BioNER. Specifically, we conduct in-distribution experiments on five biomedical datasets and out-of-distribution experiments on two datasets. Experimental results show that our method outperforms other LLMs-based methods and significantly improves the performance of BioNER. Qibin Li, Shengyuan Bai, Nai Zhou, Nianmin Yao |
BIBM | 2 |
| 2024 | Efficient Antibody Structure Refinement Using Energy-Guided SE(3) Flow MatchingabstractAntibodies are proteins produced by the immune system that recognize and bind to specific antigens, and their 3D structures are crucial for understanding their binding mechanism and designing therapeutic interventions. The specificity of antibody-antigen binding predominantly depends on the complementarity-determining regions (CDR) within antibodies.Despite recent advancements in antibody structure prediction, the quality of predicted CDRs remains suboptimal.In this paper, we develop a novel antibody structure refinement method termed FlowAB based on energy-guided flow matching. FlowAB adopts the powerful deep generative method SE(3) flow matching and simultaneously incorporates important physical prior knowledge into the flow model to guide the generation process.The extensive experiments demonstrate that FlowAB can significantly improve the antibody CDR structures. It achieves new state-of-the-art performance on the antibody structure prediction task when used in conjunction with an appropriate prior model while incurring only marginal computational overhead. This advantage makes FlowAB a practical tool in antibody engineering. Jiying Zhang, Zijing Liu, Shengyuan Bai, He Cao, Yu Li 0003, Lei Zhang 0001 |
BIBM | 3 |