VLDB 2026 Research / reviewers in the wild / expert
Ruiheng Liu
dblp:340/5478
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SafeNLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database InterfacesabstractThe rapid advancement of Large Language Models (LLMs) has driven significant progress in Natural Language Interface to Database (NLIDB). However, the widespread adoption of LLMs has raised critical privacy and security concerns. During interactions, LLMs may unintentionally expose confidential database contents or be manipulated by attackers to exfiltrate data through seemingly benign queries. While current efforts typically rely on rule-based heuristics or LLM agents to mitigate this leakage risk, these methods still struggle with complex inference-based attacks, suffer from high false positive rates, and often compromise the reliability of SQL queries. To address these challenges, we propose SafeNLIDB, a novel privacy-security alignment framework for LLM-based NLIDB. The framework features an automated pipeline that generates hybrid chain-of-thought interaction data from scratch, seamlessly combining explicit security reasoning with SQL generation. Additionally, we introduce reasoning warm-up and alternating preference optimization to overcome the multi-preference oscillations of Direct Preference Optimization (DPO), enabling LLMs to produce security-aware SQL through fine-grained reasoning without the need for human-annotated preference data. Extensive experiments demonstrate that our method outperforms both larger-scale LLMs and ideal-setting baselines, achieving significant security improvements while preserving high utility. Ruiheng Liu, Xiaobing Chen, Qiongwen Zhang, Yu Zhang 0030, Bailong Yang |
AAAI | 1 |
| 2026 | iRUC: Reducing Inter-Microservice Data Communication in Data-Intensive Systems via Unified ComputationabstractIn data-intensive microservice-based systems, frequent and large-scale inter-service communication poses a critical performance bottleneck, degrading throughput and escalating latency. Existing solutions exhibit notable limitations: microservice merging sacrifices loose coupling and system evolvability; resource-aware scheduling enhances communication efficiency but fails to reduce data volume; dynamic deployment reduces network distances while introducing compute-resource contention; and unnecessary data transfer elimination remains ineffective under massive data loads. Hence, to overcome these challenges, we propose iRUC, an approach forinter-service data communicationReduction viaUnifiedComputation. In particular, we first designGraphQL+, an executable declarative language that extends GraphQL with service-interaction semantics to achieve unified, cross-language modeling of data processing and transmission across microservices. Second, we develop anLLM-based multi-agent systemleveraging Claude 4.5 Sonnet and Gemini 2.5 Pro to automatically parse microservice code and synthesize corresponding GraphQL+ models. Third, we implement the unifiedexecution enginefor GraphQL+ models, including the database gateway that preserves microservice database autonomy while enabling cross-database queries. This design enables iRUC to perform the unified modeling and execution of data processing and transmission across microservices, thereby significantly reducing inter-service data transfer while maintaining microservice independence. Experimental evaluation on nine GitHub open-source microservice projects deployed on Huawei Cloud demonstrates iRUC’s effectiveness: compared with the unnecessary transfer elimination, dynamic deployment, and serverless computing approaches, iRUC improves throughput by 5.57×, 1.52×, and 1.87×, respectively, while reducing latency to 7.7%, 40.7%, and 37.4% of those approaches. These results show that iRUC achieves significant performance improvements in large-scale data processing scenarios. Puwei Wang, Ruiheng Liu, Keman Huang, Xiaoyong Du 0001 |
IEEE Trans. Software Eng. | 2 |
| 2025 | Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data ReplayabstractContinual Semantic Parsing (CSP) aims to train parsers to convert natural language questions into SQL across tasks with limited annotated examples, adapting to dynamically updated databases in real-world scenarios. Previous studies mitigate this challenge by replaying historical data or employing parameter-efficient tuning (PET), but they often violate data privacy or rely on ideal continual learning settings. To address these issues, we propose a new Large Language Model (LLM)-Enhanced Continuous Semantic Parsing method, named LECSP, which alleviates forgetting while encouraging generalization, without requiring real data replay or ideal settings. Specifically, it first analyzes the commonalities and differences between tasks from the SQL syntax perspective to guide LLMs in reconstructing key memories and improving memory accuracy through calibration. Then, it uses a task-aware dual-teacher distillation framework to promote the accumulation and transfer of knowledge during sequential training. Experimental results on two CSP benchmarks show that our method significantly outperforms existing methods, even those utilizing data replay or ideal settings. Additionally, we achieve generalization performance beyond upper limits, better adapting to unseen tasks. Ruiheng Liu, Yanqi Song, Yu Zhang 0030, Bailong Yang |
AAAI | 1 |
| 2025 | Discarding the Crutches: Adaptive Parameter-Efficient Expert Meta-Learning for Continual Semantic ParsingabstractContinual Semantic Parsing (CSP) enables parsers to generate SQL from natural language questions in task streams, using minimal annotated data to handle dynamically evolving databases in real-world scenarios. Previous works often rely on replaying historical data, which poses privacy concerns. Recently, replay-free continual learning methods based on Parameter-Efficient Tuning (PET) have gained widespread attention. However, they often rely on ideal settings and initial task data, sacrificing the model’s generalization ability, which limits their applicability in real-world scenarios. To address this, we propose a novel Adaptive PET eXpert meta-learning (APEX) approach for CSP. First, SQL syntax guides the LLM to assist experts in adaptively warming up, ensuring better model initialization. Then, a dynamically expanding expert pool stores knowledge and explores the relationship between experts and instances. Finally, a selection/fusion inference strategy based on sample historical visibility promotes expert collaboration. Experiments on two CSP benchmarks show that our method achieves superior performance without data replay or ideal settings, effectively handling cold start scenarios and generalizing to unseen tasks, even surpassing performance upper bounds. Ruiheng Liu, Yanqi Song, Yu Zhang 0030, Bailong Yang |
COLING | 1 |
| 2025 | Graph-Embedded Structure-Aware Perceptual Hashing for Neural Network Protection and Piracy DetectionabstractThe advancement of AI technology has significantly influenced production activities, increasing the focus on copyright protection for AI models. The perceptual hashing of the model offers an efficient solution for retrieve the pirated models. Existing methods, such as handcrafted feature-based and dual-branch network-based perceptual hashing, have proven effective in detecting pirated models. However, these approaches often struggle to differentiate nonpirated models, leading to frequent false positives in model authentication and protection. To address this challenge, this paper proposes a structurally-aware perceptual model hashing technique that achieved reduced false positives while maintaining high true positive rates. Specifically, we introduce a method for converting the diverse neural network structures into graph structures suitable for DNN processing, then utilize a graph neural network to learn their structural features representation. Our approach integrates perceptual parameter-based model hashing, achieving robust performance with higher detection accuracy and fewer false positives. The experimental results show that the proposed method has only 3% false alarm rate when detecting the non-pirated model, and the detection accuracy of the pirated model reaches more than 98%. Ruiheng Liu, Boyao Zhao, Kejiang Chen, Weiming Zhang 0001 |
CVPR | 1 |
| 2025 | UVS: A Novel Underwater Vehicle with Integrated VCMS-Thrusters Hybrid Architecture for Enhanced Attitude RegulationabstractAutonomous Underwater Vehicles (AUVs) require energy-efficient and responsive attitude control for underwater operations. We present UVS, a novel underwater vehicle that combines Variable Center of Mass System (VCMS) and thrusters for hybrid attitude regulation. Through multi-objective optimization of the VCMS structure, we achieved a 5.19% larger pitch angle range while reducing space occupation by 15.72%. Pool experiments demonstrated near-linear pitch control from 17.5° to 172.5° with stable horizontal-vertical mode transitions. Our proposed collaborative control method integrates VCMS and thruster advantages, enabling rapid convergence to target attitudes with long-term stability. The results show UVS’s potential for energy-efficient, wide-range attitude control in mobile ocean sensing applications. Suohang Zhang, Shipang Qian, Ruiheng Liu, Xinyu Fei |
IROS | 3 |
| 2025 | GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic SecurityabstractMinimum distortion steganography is currently the mainstream method for modification-based steganography. A key issue in this method is how to define steganographic distortion. With the rapid development of deep learning technology, the definition of distortion has evolved from manual design to deep learning design. Concurrently, rapid advancements in image generation have made generated images viable as cover media. However, existing distortion design methods based on machine learning do not fully leverage the advantages of generated cover media, resulting in suboptimal security performance. To address this issue, we propose GIFDL (Generated Image Fluctuation Distortion Learning), a steganographic distortion learning method based on the fluctuations in generated images. Inspired by the idea of natural steganography, we take a series of highly similar fluctuation images as the input to the steganographic distortion generator and introduce a new GAN training strategy to disguise stego images as fluctuation images. Experimental results demonstrate that GIFDL, compared with state-of-the-art GAN-based distortion learning methods, exhibits superior resistance to steganalysis, increasing the detection error rates by an average of 3.30% across three steganalyzers. Xiangkun Wang, Kejiang Chen, Yuang Qi, Ruiheng Liu, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Mitigating the Data Communication Overhead in Microservice-based Data-intensive SystemsabstractMicroservice architecture is favored for its loose coupling, reusability, and scalability. However, in data-intensive systems where data is the primary and permanent assets, the dynamic inter-microservice communication leads to a large amount of inter-microservice data transfer. This leads to a significant reduction in throughput and an increase in latency. While existing strategies, including microservice decomposition, deployment and resource optimization, show promise, they overlook the unnecessary inter-microservice data communication in practice, which includes the excessive data exposure and carryover data. This motivates us to develop iRUC, an integrated approach to remove the unnecessary inter-microservice data communications, through integrating the Excessive Data Transmission Removal, Carryover Microservice Upgrade, and Query Language (QL) Statement Composition. By implementing a microservice based data intensive system and deploying it on the public cloud environment, the experimental results confirm the effectiveness of our approach, achieving on average 4.2 × to 5.3 × throughput improvement and 78.6% to 85.9% latency reduction. Puwei Wang, Ruiheng Liu, Bo Liu 0010, Keman Huang, Xiaoyong Du 0001 |
ICWS | 2 |
| 2024 | Robust and resource-efficient table-based fact verification through multi-aspect adversarial contrastive learning
Ruiheng Liu, Yu Zhang 0030, Bailong Yang, Qi Shi 0002, Luogeng Tian |
Inf. Process. Manag. | 1 |