VLDB 2026 Research / reviewers in the wild / expert
Ruihan Hu
dblp:234/7554
· DBLP profile ↗
19ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0001-8525-2503ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning ModelsabstractLarge Reasoning Models (LRMs) have rapidly gained prominence for their strong performance in solving complex tasks. Many modern black-box LRMs expose the intermediate reasoning traces through APIs to improve transparency (e.g., Gemini-2.5 and Claude-sonnet). Despite their benefits, we find that these traces can leak membership signals, creating a new privacy threat even without access to token logits used in prior attacks. In this work, we initiate the first systematic exploration of Membership Inference Attacks (MIAs) on black-box LRMs. Our preliminary analysis shows that LRMs produce confident, recall-like reasoning traces on familiar training member samples but more hesitant, inference-like reasoning traces on non-members. The representations of these traces are continuously distributed in the semantic latent space, spanning from familiar to unfamiliar samples. Building on this observation, we propose BlackSpectrum, the first membership inference attack framework targeting the black-box LRMs. The key idea is to construct a recall–inference axis in the semantic latent space, based on representations derived from the exposed traces. By locating where a query sample falls along this axis, the attacker can obtain a membership score and predict how likely it is to be a member of the training data. Additionally, to address the limitations of outdated datasets unsuited to modern LRMs, we provide two new datasets to support future research, arXivReasoning and BookReasoning. Empirically, exposing reasoning traces greatly increases the vulnerability of LRMs to MIAs, boosting attack accuracy by up to 23.8%, AUC by 29.9%, and nearly doubling TPR@5%FPR. Our findings highlight the need for LRM companies to balance transparency in intermediate reasoning traces with privacy preservation. Ruihan Hu, Yuming Shang, Wei Luo 0016, Xi Zhang 0008 |
WWW | 1 |
| 2026 | Joint Optimization of Fine-Grained Representation and Workflow Orchestration in Metaverse Articulated Manipulation Auto-Generation by VLA MethodabstractArticulated manipulation represents a nuanced form of trajectory generation, offering a series of service APIs to guide robot arms to understand the structure of articulated objects. The trajectory of the articulated products in the real world is expected to be guided by the complex steps of robotic arms in the metaverse environment. However, the generated joint torque intervals are varied when it comes to applying to different types of robots and appearances of objects. As a result, these intervals exhibit limitations in category-agnostic articulated manipulation. To address these issues, this study proposes a simple articulated manipulation auto-generation method via joint optimization of Vision-Language-Action model (AMAG-JOVLA), which has geometric-centric, category-agnostic, kinematics-aware characteristics. AMAG-JOVLA provides Geometric-Centric Multi-instance Multi-label (GCMIML) representation for articulated manipulation as a multi-instance multi-label problem. Moreover, to generate motion waypoints with fewer spatial, temporal, curvature, and direction errors, a kinematic-aware prompting method, the Geometric Thought Planner (GTP), it is used to build the adaptable manipulation workflows within the inference process of articulated manipulation. GTP also contains a joint optimization mechanism called GTP-Search, which accelerates updating of the label set of GCMIML as a tree-based search problem. The comparative, ablation, and category-agnostic experiments conducted in the metaverse and the real-world environments indicate that the performance of AMAG-JOVLA outperforms that of other comparative methods. Ruihan Hu, Xiangdong He, Feiyang Huang, Jiaxing Zhao, Xinrui Cheng, Zhongjie Wang 0003 |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Automated Detection of Pre-training Text in Black-box LLMsabstractDetecting whether a given text is a member in the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model parameters or token probabilities), making them ineffective in the black-box setting, where only input and output texts are accessible. Although some methods have been proposed for the black-box setting, they rely on massive manual efforts such as designing complicated questions or instructions. To address these issues, we propose VeilProbe, the first framework for automatically detecting LLMs' pre-training texts in a black-box setting without human intervention. VeilProbe utilizes a sequence-to-sequence mapping model to infer the latent mapping feature between the input text and the corresponding output suffix generated by the LLM. Then it performs the key token perturbations to obtain more distinguishable membership features. Additionally, considering real-world scenarios where the ground-truth training text samples are limited, a prototype-based membership classifier is introduced to alleviate the overfitting issue. Extensive evaluations on three widely used datasets demonstrate that our framework is effective and superior in the black-box setting. Ruihan Hu, Yuming Shang, Jiankun Peng, Wei Luo 0016, Yazhe Wang, Xi Zhang 0008 |
IJCAI | 1 |
| 2025 | SGGAA: a robust adversarial generation based on synthesizing point cloud and mesh space
Ruihan Hu |
Neural Comput. Appl. | 1 |
| 2025 | Device Selection and Resource Allocation With Semi-Supervised Method for Federated Edge LearningabstractWith the rapid growth of distributed learning and workflow orchestration, Federated Edge Learning has emerged as a solution, enabling multiple edge devices to collaboratively train a large model without the need for sharing raw data. Beyond considering bandwidth and computational resource limitations in the Internet of Things (IoT) environment, it is crucial to address the issue of IoT devices often collecting data that lacks timely annotations, which can lead to latency and label deficiency issues. In most Federated Edge Learning mechanisms, clients’ weights are selected for offloading to the server. In this paper, we propose a solution for dynamic edge selection and wireless network allocation under semi-supervised and privacy protection settings, termed Semi-supervised Scheduling and Allocation Optimization for Federated Edge Learning (SSAFL). SSAFL is designed to adapt to various scenarios, including channel state variations, device heterogeneity, resource incentives, deadline control, label deficiencies, and Non-IID data distributions. This adaptability is achieved through the utilization of an Incentive Optimization framework that encompasses bandwidth allocation and device scheduling policies. Within SSAFL, we introduce the concept of a weighted bipartite graph network to tackle the Incentive Optimization problem and achieve a balance in large-scale optimization of device selection. Additionally, to address the label deficiency issue, we devise a Dynamic Timer for deadline control for each client. Comprehensive and confidential results demonstrate that our proposed approach significantly outperforms other Federated Edge Learning baselines. Ruihan Hu, Haochen Yuan 0001, Daimin Tan, Zhongjie Wang 0003 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | GCPN: A Group Connected based Method for Continual Vertical Federated Recommender Systems in Data EcosystemsabstractData ecosystems (DE) are the future directions of data management and play a vital role in unlocking the value of data. Service Recommender Systems (RS) are typical applications in DEs. For example, deep learning-based RS on the basis of extensive data from various fields can help organizations obtain valuable insights of data. As organizations from different fields share data for better recommendation services, the risk of privacy leakage which is harmful to DEs increases. Due to privacy concerns, Vertical Federated Learning (VFL), a privacy-preserving computing technology for joint learning and privacy recommendation models among organizations in various fields, has garnered significant attention. Existing VFL methods are training models with static data from specific fields when there are new recommendation scenarios or fields. In addition, the models are fixed after one training session. Therefore, these models can only be applied to several specific recommendation fields and they can’t utilize continuously generated data that corresponds to various fields, posing challenges for long-term and extensive cooperation. To tackle these challenges, we introduce Vertical Federated Continual Learning (VFCL), which extends Continual Learning (CL) into the VFL framework to enable VFL models to sustainably adapt to new scenarios. We discuss feasible solutions based on existing CL methods. Furthermore, we propose GCPN, a method based on a dynamic architecture in VFCL. GCPN introduces fewer parameters for each new field by utilizing group connected layers and scale layers, eliminating the need for storing or using past data. It effectively alleviates the problem of catastrophic forgetting, a major issue in CL, while preserving privacy in joint recommendations. To evaluate GCPN, we construct the VFCL scenario using Amazon's public recommendation datasets. Experiments demonstrate that our method enhances the effectiveness of most tasks in CL and VFCL scenarios. Haochen Yuan 0001, Xiang He 0002, Ruihan Hu, Zhongjie Wang 0003, Yunqing Feng, Lecheng Gong |
ICWS | 3 |
| 2024 | MDSSN: An end-to-end deep network on triangle mesh parameterization
Ruihan Hu, Zhongjie Wang 0003 |
Knowl. Based Syst. | 1 |
| 2022 | RDC-SAL: Refine distance compensating with quantum scale-aware learning for crowd counting and localization
Ruihan Hu, Qi Wu 0003, Qinglong Mo, Jingbin Li |
Appl. Intell. | 1 |
| 2022 | Multi-expert learning for fusion of pedestrian detection bounding box
Ruihan Hu, Zhao-Hui Sun, Ming Li 0055 |
Knowl. Based Syst. | 2 |
| 2022 | Nonparametric Hierarchical Hidden Semi-Markov Model for Brain Fatigue Behavior Detection of Pilots During FlightabstractThe evaluation of pilot brain activity is very important for flight safety. This study proposes a Hidden semi-Markov Model with Hierarchical prior to detect brain activity under different flight tasks. A dynamic student mixture model is proposed to detect the outlier of emission probability of HSMM. Instantaneous spectrum features are also extracted from EEG signals. Compared with other latent variable models, the proposed model shows excellent performance for the automatic inference of brain cognitive activity of pilots. The results indicate that the consideration of hierarchical model and the emission probability with${t}$mixture model improves the recognition performance for Pilots’ fatigue cognitive level. Qi Wu 0003, Limin Zhu 0001, Gui-Jiang Li, Ruihan Hu, Gui-Rong Zhou |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Inferring Flight Performance Under Different Maneuvers With Pilot's Multi-Physiological ParametersabstractThe relationship between flight performance and multi-physiological parameters under different flight operating patterns is unknown. This work proposes a Stacked Gaussian Process Network (SGPN) to reveal it. SGPN is a multi-layer network model formed by recursion from a regular Gaussian process and random disturbance. This work constructs an auxiliary variable strategy with the induced points to improve its learning efficiency, thus leading to a sparse SGPN model. In it, a Gaussian process acts as an activation function of each node, but the entire model is no longer a Gaussian process and thus very challenging to solve it. This work presents its solution via variational approximate inference. Experimental results of pilot flight performance evaluation show that the proposed model has stronger learning and generalization ability than its seven competitive peers. It is able to approximate non-linear coupling relationship between multi-physiological parameters and flight height differences. Qi Wu 0003, MengChu Zhou, Pengwen Xiong, Ruihan Hu, Yu-Wen Jie |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Non-spike timing-dependent plasticity learning mechanism for memristive neural networks
Zhihua Wang 0002, Ruihan Hu, Qi Wu 0003 |
Appl. Intell. | 4 |
| 2021 | Ensemble echo network with deep architecture for time-series modeling
Ruihan Hu, Qi Wu 0003, Sheng Chang 0003 |
Neural Comput. Appl. | 1 |
| 2021 | DMMAN: A two-stage audio-visual fusion framework for sound separation and event localization
Ruihan Hu, Songbin Zhou, Sheng Chang 0003, Qijun Huang, Yisen Liu, Qi Wu 0003 |
Neural Networks | 1 |
| 2020 | Fully memristive spiking-neuron learning framework and its applications on pattern recognition and edge detection
Shizhuo Ye, Ruihan Hu, Hao Wang 0046, Jin He 0002, Qijun Huang, Sheng Chang 0003 |
Neurocomputing | 4 |
| 2019 | The MBPEP: a deep ensemble pruning algorithm providing high quality uncertainty prediction
Ruihan Hu, Qijun Huang, Sheng Chang 0003, Hao Wang 0046, Jin He 0002 |
Appl. Intell. | 1 |
| 2019 | Monitor-Based Spiking Recurrent Network for the Representation of Complex Dynamic PatternsabstractNeural networks are powerful computation tools for mimicking the human brain to solve realistic problems. Since spiking neural networks are a type of brain-inspired network, called the novel spiking system, Monitor-based Spiking Recurrent network (MbSRN), is derived to learn and represent patterns in this paper. This network provides a computational framework for memorizing the targets using a simple dynamic model that maintains biological plasticity. Based on a recurrent reservoir, the MbSRN presents a mechanism called a 'monitor' to track the components of the state space in the training stage online and to self-sustain the complex dynamics in the testing stage. The network firing spikes are optimized to represent the target dynamics according to the accumulation of the membrane potentials of the units. Stability analysis of the monitor conducted by limiting the coefficient penalty in the loss function verifies that our network has good anti-interference performance under neuron loss and noise. The results of solving some realistic tasks show that the MbSRN not only achieves a high goodness-of-fit of the target patterns but also maintains good spiking efficiency and storage capacity. Ruihan Hu, Qijun Huang, Hao Wang 0046, Jin He 0002, Sheng Chang 0003 |
Int. J. Neural Syst. | 1 |
| 2019 | Design of high-resolution quantization scheme with exp-Golomb code applied to compression of special images
Qijun Huang, Ruihan Hu |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Efficient Multispike Learning for Spiking Neural Networks Using Probability-Modulated Timing MethodabstractError functions are normally based on the distance between output spikes and target spikes in supervised learning algorithms for spiking neural networks (SNNs). Due to the discontinuous nature of the internal state of spiking neuron, it is challenging to ensure that the number of output spikes and target spikes kept identical in multispike learning. This problem is conventionally dealt with by using the smaller of the number of desired spikes and that of actual output spikes in learning. However, if this approach is used, information is lost as some spikes are neglected. In this paper, a probability-modulated timing mechanism is built on the stochastic neurons, where the discontinuous spike patterns are converted to the likelihood of generating the desired output spike trains. By applying this mechanism to a probability-modulated spiking classifier, a probability-modulated SNN (PMSNN) is constructed. In its multilayer and multispike learning structure, more inputs are incorporated and mapped to the target spike trains. A clustering rule connection mechanism is also applied to a reservoir to improve the efficiency of information transmission among synapses, which can map the highly correlated inputs to the adjacent neurons. Results of comparisons between the proposed method and popular the SNN algorithms showed that the PMSNN yields higher efficiency and requires fewer parameters. Ruihan Hu, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |