VLDB 2026 Research / reviewers in the wild / expert
Bailong Yang
dblp:224/2383 · also BaiLong Yang
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data models and query languages · 68% Query processing and optimization · 32% | |
| Network and information security
1 paper |
Systems and software security · 100% | |
| Computer networks
1 paper |
Network performance modeling · 77% Internet architecture and protocols · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Energy-efficient computing · 100% |
Topics — the 4 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data models and query languages › natural language interface
natural language interface to database |
1.0 | 1 | 2026 | SafeNLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database Interfaces · AAAI 2026 |
Query processing and optimization
semantic parsing |
0.9 | 1 | 2025 | Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data Replay · AAAI 2025 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.9 | 1 | 2025 | Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data Replay · AAAI 2025 |
Energy-efficient computing
power management |
0.4 | 1 | 2020 | Towards Power Efficient High Performance Packet I/O · IEEE Trans. Parallel Distributed Syst. 2020 |
Methods — techniques the papers use, named apart from their topics
direct preference optimization · 2.0chain-of-thought data generation · 2.0alternating preference optimization · 2.0parameter-efficient tuning · 0.9large language model · 0.9knowledge distillation · 0.9traffic-aware power conservation · 0.9pause instruction · 0.9LPI · 0.9DVFS · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SafeNLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database InterfacesabstractThe rapid advancement of Large Language Models (LLMs) has driven significant progress in Natural Language Interface to Database (NLIDB). However, the widespread adoption of LLMs has raised critical privacy and security concerns. During interactions, LLMs may unintentionally expose confidential database contents or be manipulated by attackers to exfiltrate data through seemingly benign queries. While current efforts typically rely on rule-based heuristics or LLM agents to mitigate this leakage risk, these methods still struggle with complex inference-based attacks, suffer from high false positive rates, and often compromise the reliability of SQL queries. To address these challenges, we propose SafeNLIDB, a novel privacy-security alignment framework for LLM-based NLIDB. The framework features an automated pipeline that generates hybrid chain-of-thought interaction data from scratch, seamlessly combining explicit security reasoning with SQL generation. Additionally, we introduce reasoning warm-up and alternating preference optimization to overcome the multi-preference oscillations of Direct Preference Optimization (DPO), enabling LLMs to produce security-aware SQL through fine-grained reasoning without the need for human-annotated preference data. Extensive experiments demonstrate that our method outperforms both larger-scale LLMs and ideal-setting baselines, achieving significant security improvements while preserving high utility. Ruiheng Liu, Xiaobing Chen, Qiongwen Zhang, Yu Zhang 0030, Bailong Yang |
AAAI | 6 |
| 2025 | Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data ReplayabstractContinual Semantic Parsing (CSP) aims to train parsers to convert natural language questions into SQL across tasks with limited annotated examples, adapting to dynamically updated databases in real-world scenarios. Previous studies mitigate this challenge by replaying historical data or employing parameter-efficient tuning (PET), but they often violate data privacy or rely on ideal continual learning settings. To address these issues, we propose a new Large Language Model (LLM)-Enhanced Continuous Semantic Parsing method, named LECSP, which alleviates forgetting while encouraging generalization, without requiring real data replay or ideal settings. Specifically, it first analyzes the commonalities and differences between tasks from the SQL syntax perspective to guide LLMs in reconstructing key memories and improving memory accuracy through calibration. Then, it uses a task-aware dual-teacher distillation framework to promote the accumulation and transfer of knowledge during sequential training. Experimental results on two CSP benchmarks show that our method significantly outperforms existing methods, even those utilizing data replay or ideal settings. Additionally, we achieve generalization performance beyond upper limits, better adapting to unseen tasks. Ruiheng Liu, Yanqi Song, Yu Zhang 0030, Bailong Yang |
AAAI | 5 |
| 2025 | Discarding the Crutches: Adaptive Parameter-Efficient Expert Meta-Learning for Continual Semantic ParsingabstractContinual Semantic Parsing (CSP) enables parsers to generate SQL from natural language questions in task streams, using minimal annotated data to handle dynamically evolving databases in real-world scenarios. Previous works often rely on replaying historical data, which poses privacy concerns. Recently, replay-free continual learning methods based on Parameter-Efficient Tuning (PET) have gained widespread attention. However, they often rely on ideal settings and initial task data, sacrificing the model’s generalization ability, which limits their applicability in real-world scenarios. To address this, we propose a novel Adaptive PET eXpert meta-learning (APEX) approach for CSP. First, SQL syntax guides the LLM to assist experts in adaptively warming up, ensuring better model initialization. Then, a dynamically expanding expert pool stores knowledge and explores the relationship between experts and instances. Finally, a selection/fusion inference strategy based on sample historical visibility promotes expert collaboration. Experiments on two CSP benchmarks show that our method achieves superior performance without data replay or ideal settings, effectively handling cold start scenarios and generalizing to unseen tasks, even surpassing performance upper bounds. Ruiheng Liu, Yanqi Song, Yu Zhang 0030, Bailong Yang |
COLING | 5 |
| 2024 | Robust and resource-efficient table-based fact verification through multi-aspect adversarial contrastive learning
Ruiheng Liu, Yu Zhang 0030, Bailong Yang, Qi Shi 0002, Luogeng Tian |
Inf. Process. Manag. | 3 |
| 2020 | Limited-Budget Output Consensus for Descriptor Multiagent Systems With Energy ConstraintsabstractThis article deals with limited-budget output consensus for descriptor multiagent systems with two types of switching communication topologies, that is, switching connected ones and jointly connected ones. First, a singular dynamic output feedback control protocol with switching communication topologies is proposed on the basis of the observable decomposition, where an energy constraint is involved and protocol states of neighboring agents are utilized to derive a new two-step design approach of gain matrices. Then, limited-budget output consensus problems are transformed into asymptotic stability ones and a valid candidate of the output consensus function is determined. Furthermore, sufficient conditions for limited-budget output consensus design and analysis for two types of switching communication topologies are proposed, respectively, and an explicit expression of the output consensus function is given, which is identical for two types of switching communication topologies. Finally, two numerical simulations are shown to demonstrate theoretical conclusions. Jianxiang Xi, Bailong Yang |
IEEE Trans. Cybern. | 4 |
| 2020 | Towards Power Efficient High Performance Packet I/OabstractRecently, high performance packet I/O frameworks continue to flourish for their ability to process packets from high-speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. However, exploiting power management techniques such as DVFS and LPI in the frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., actual CPU utilization, available idle period, or the target frequency) required by these techniques. In this article, we establish a model that can formulate the packet processing flow of high performance packet I/O to help and address the above challenges. From the model, we can deduce the information needed for power management techniques, and gain the insights to balance the power and latency. After suggesting to use pause instruction to reduce CPU power within short idle period, we propose two approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction with little latency increase. Wenxue Cheng, Tong Zhang 0018, Fengyuan Ren, Bailong Yang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Power Efficient High Performance Packet I/OabstractRecently, high performance packet I/O frameworks are expected an extensive application for their ability to process packets from 10Gbps or higher speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling technique. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. Meanwhile, exploiting power management techniques such as DVFS and LPI in high performance packet I/O frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., the actual CPU utilization, available idle period, or the target frequency) required by power management techniques. In this paper, we establish an analytical model that can formulate the packet processing flow of high performance packet I/O to help address the above challenges. From the analytical model, we can deduce the actual CPU utilization and average idle period in different traffic load, and gain the insight to choose CPU frequency that can appropriately balance the power consumption and packet latency. Then, we propose two simple but effective approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction (35.90% and 34.43% on average respectively) while incurring < 1 μs of latency increase. Wenxue Cheng, Tong Zhang 0018, Jing Xie 0005, Fengyuan Ren, Bailong Yang |
ICPP | 6 |
| 2015 | Consensus transformation for multi-agent systems with topology variances and time-varying delays
Jianxiang Xi, Jinying Wu, Bailong Yang |
Neurocomputing | 4 |
| 2015 | Stable-protocol admissible synchronizability for high-order singular complex networks with switching topologies
Jianxiang Xi, Bailong Yang |
Inf. Sci. | 4 |