VLDB 2026 Research / reviewers in the wild / expert
Yulun Song
dblp:153/9941
· DBLP profile ↗
12ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0006-1029-7823ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyDRA: Hyperbolic dual-geometry representation alignment for knowledge-aware recommendation
Shaoxing Zhang, Shuai Zhao 0001, Yulun Song |
Neurocomputing | 4 |
| 2026 | zkVFL: Verifiable Federated Learning for Free-Rider Attacks via Efficient Zero-Knowledge ProofsabstractFederated Learning (FL) enables model training on distributed devices while preserving data privacy. However, malicious clients can submit fabricated model updates to fraudulently obtain training rewards, a behavior known as free-rider attacks. Existing detection-based solutions analyze anomalies in model updates but lack direct evidence of local training, making it fail to fully prevent free-riders. To address this limitation, we propose zkVFL, a verifiable FL framework leveraging Zero-Knowledge Proofs (ZKP) to ensure the integrity of local training while preserving privacy. To reduce the computational overhead of proof generation in ZKP, zkVFL introduces two novel techniques: (i) anomaly-aware client sampling to selectively perform ZKP verification and (ii) A recursive ZKP protocol (ReMPoT), incorporating a pruning-based layer selection technique, reduces proof generation costs. Experimental results demonstrate that zkVFL improves the accuracy and convergence of FL training under free-rider attacks while significantly reducing the computational and memory overhead of proof generation on resource-constrained devices. Tianyu Kang, Di Wu 0065, Yulun Song, Yunlong Xie, Li Guo 0004 |
IEEE Internet Things J. | 5 |
| 2026 | DMS-P$^{2}$2CQ: Privacy-Preserving Collaborative Query Protocol for Distributed Multi-Server SystemsabstractA privacy-preserving collaborative query protocol for distributed multi-server systems (DMS-P$^{2}$CQ) enables a querying party to interact with multiple independent servers using an identifier$x_{u}$and receive a categorical decision (e.g.,Good/Moderate/Poor) determined by the total number of servers whose datasets contain$x_{u}$. This setting is motivated by financial applications such as credit assessment, where the querying party must not reveal$x_{u}$to data-owning institutions, and each institution must protect its proprietary user list. To meet both the functionality and privacy requirements of the querying party and the servers, we present two protocols that represent a privacy-efficiency trade-off. DMS-P$^{2}$CQ$_{1}$is a lightweight protocol inspired by OPRF-based PSI. It achieves identifier privacy and hides the identifier-to-server membership relation from the querying party via a two-stage OPRF with an aggregation/re-randomization server. However, it reveals the aggregate count to a randomly selected leader server and relies on a non-collusion assumption between two special servers. DMS-P$^{2}$CQ$_{2}$strengthens privacy by secret-sharing the count so that no single party learns the true aggregate count and removes the need for a trusted re-randomization server. We prove the security of two protocols under the semi-honest model. Experimental results on our local testbed show that with 20 servers and a dataset size of$2^{20}$, the runtimes for DMS-P$^{2}$CQ$_{1}$and DMS-P$^{2}$CQ$_{2}$are approximately 8.034s and 16.697s, respectively. Huihui Zhu 0001, Fei Tang 0001, Jinyong Shan, Ping Wang 0086, Yulun Song, Yunlong Xie |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Reliable Active Apprenticeship LearningabstractWe propose a learning problem, which we call reliable active apprenticeship learning, for which we define a learning algorithm providing optimal performance guarantees, which we further show are sharply characterized by the eluder dimension of a policy class. In this setting, a learning algorithm is tasked with behaving optimally in an unknown environment given by a Markov decision process. The correct actions are specified by an unknown optimal policy in a given policy class. The learner initially does not know the optimal policy, but it has the ability to query an expert, which returns the optimal action for the current state. A learner is said to be reliable if, whenever it takes an action without querying the expert, its action is guaranteed to be optimal. We are then interested in designing a reliable learner which does not query the expert too often. We propose a reliable learning algorithm which provably makes the minimal possible number of queries, which we show is precisely characterized by the eluder dimension of the policy class. We further extend this to allow for imperfect experts, modeled as an oracle with noisy responses. We study two variants of this, inspired by noise conditions from classification: namely, Massart noise and Tsybakov noise. In both cases, we propose a reliable learning strategy which achieves a nearly-minimal number of queries, and prove upper and lower bounds on the optimal number of queries in terms of the noise conditions and the eluder dimension of the policy class. Steve Hanneke, Liu Yang 0001, Gongju Wang, Yulun Song |
ALT | 4 |
| 2025 | TrustMonitor: A Secured Container Framework via Virtualization and Pointer Authentication for Cross-Domain Data SharingabstractWe propose TrustMonitor, a lightweight runtime container protection framework designed for secure data sharing in cross-domain scenarios. TrustMonitor establishes a trusted container execution environment through the coordinated use of memory isolation, authenticated memory access, and identity binding. It leverages Extended Page Tables (EPT) to enforce hardware-level separation between containers and the host OS, forming the foundation for isolating sensitive data from untrusted system components. Building upon this isolation, TrustMonitor integrates software-emulated Pointer Authentication Codes (PAC) to authorize memory access on a per-context basis, ensuring that only verified execution paths can interact with protected container memory. To complete the trust chain, each container is bound to a virtual Trusted Platform Module (vTPM), which performs integrity measurements during instantiation and safeguards runtime secrets. Collectively, they establish a trusted execution foundation that mitigates kernel-level threats and runtime compromise, thereby securing sensitive data throughout its lifecycle in cross-domain sharing environments. Experimental results demonstrate that TrustMonitor effectively safeguards the confidentiality and integrity of shared data. Furthermore, it achieves these security guarantees with less than 15% computational overhead and no more than 25% latency increase, making it a practical and efficient solution for secure cross-domain data sharing environments. Yulun Song, Yunlong Xie |
TrustCom | 2 |
| 2024 | NAGG: Noised graph node feature aggregations for preserving privacyabstractGraph node feature aggregation is a very common basic operation in graph data query and graph neural networks. We design three noised aggregation methods to protect the node features of the graph and analyze the privacy of these methods from the perspective of differential privacy. NAGG1, our first method, adds Gaussian noise to the aggregation result, extending the classic Gaussian mechanism of differential privacy directly to the graph domain. Subsequently, we design NAGG2 which adds noise to the raw data of each sample and then performs the aggregation. This makes it possible to decouple the data owner and subsequent data users. We impose some restrictions on the original definition of differential privacy, and under these new assumptions, we demonstrate that the effects of Gaussian noise in NAGG2 can be achieved with any other type of noise that meets the requirements. We call this method NAGG3, which significantly enhances the flexibility of added noise in theory. For these methods, we instantiate them with mean aggregation and graph convolutional aggregation. We analyze their privacy protection using the framework of differential privacy and evaluate their utility through experiments. Yinghao Song, Mingjian Ni, Shengzhong Tan, Dazhong Li, Huiting Zhao, Yulun Song |
TrustCom | 8 |
| 2024 | Enhancing trust and privacy in distributed networks: a comprehensive survey on blockchain-based federated learning
Ji Liu 0003, Chunlu Chen, Yulun Song, Jingbo Zhou 0003, Bo Jing, Dejing Dou |
Knowl. Inf. Syst. | 5 |
| 2023 | Pessimistic Adversarially Regularized Learning for Graph Embedding
Yinghao Song, Hanbin Feng, Yulun Song, Gongju Wang |
ADMA (3) | 5 |
| 2023 | PAGAE: Improving Graph Autoencoder by Dual Enhanced Adversary
Gongju Wang, Hanbin Feng, Yulun Song, Yinghao Song |
CogSci | 5 |
| 2023 | D-AE: A Discriminant Encode-Decode Nets for Data Generation
Gongju Wang, Yulun Song, Mingjian Ni, Quanda Wang, Xingru Huang |
CollaborateCom (2) | 2 |
| 2017 | A method for extracting vegetation information of Urban underlaying surface oriented to ECO-environmental quality assessmentabstractIn the traditional classification method, for the pixels, which the proportion of vegetation distribution is small, that is, the area occupied by vegetation in the pixel is less than half a pixel, they are often considered as a non-vegetation area. Such problems, in the highly developed cities, are the most common. Therefore, in order to more accurately compute the highly developed, especially hardened urban underlying surface, in this study, a new method for estimating vegetation area was proposed. The main urban area of Beijing was chosen as the study area, and the 1km grid was used as the evaluation unit. Finally, we calculated the urban development intensity index and analyzed the effectiveness of the new method of vegetation area extraction in practical application. The results show that, compared with the traditional classification methods, the index model can significantly improve the accuracy of urban vegetation information extraction and can effectively extract the information of urban underlying surface vegetation information, reflecting the intensity of urban development, and provide technical support for ecological environmental quality assessment. Yulun Song |
IGARSS | 2 |
| 2014 | Development of a Universal Platform for Hardware In-the-Loop Testing of MicrogridsabstractThe operation of a microgrid becomes significantly complex with the high penetration of distributed energy resources (DERs), demand-side management, market operation, and disconnection and reconnection to the utility grid. Therefore, development of advanced tools/platforms for testing operation and control of microgrid has attracted more and more attention nowadays. The current literature reveals that the microgrid's control and management are designed to be tested either in a numerical simulation approach, or only under a specific hardware device/experiment environment; they do not deal with a comprehensive platform capable of easily executing very complex applications built by composing required functionalities in a standardized, easy-to-use, and well-defined way. To address the problem of limited testing functions of existing tools/platforms, a hardware-in-the-loop (HIL) approach, in particular combining a power-HIL (PHIL) and a signal-HIL (SHIL), is proposed in this paper. Such an approach is suitable for testing the system-level controller energy management systems (EMSs) and hardware controllers at signal level, as well as hardware devices like power converters at power level. Hence, this platform is designed for flexibility and universality. The HIL platform is presented in this work and its performance is demonstrated in a sample application. Yulun Song, Ji Guo, Antonello Monti |
IEEE Trans. Ind. Informatics | 2 |