VLDB 2026 Research / reviewers in the wild / expert
Xinran Zheng
dblp:313/6537
· DBLP profile ↗
20ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 2 first-author · 9 since 2021Security and privacy · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KNOWML: Improving Generalization of ML-NIDS with Attack Knowledge GraphsabstractAnomaly-based ML-NIDS (A-NIDS) model normal network behavior from benign data and classify deviations from this baseline as anomalies, theoretically enabling the detection of evolving attack variants without labeled attack data. The ability of A-NIDS to generalize critically depends on the quality of the feature space representing network behavior. However, the requirement for feature spaces that encode attack-relevant semantics has received little attention and remains poorly understood. As a consequence, these systems still struggle to meet practical operational constraints (low false positive rates without compromising detection performance and generalization to attack variants). We identify two limitations in the current feature spaces. First, Out-of-Dimension Blindness, where features do not capture essential attack mechanism properties. Second, Attack Strategy Aggregation Failure, where features cannot encode composite attack behaviors. Moreover, we demonstrate that two SotA data-driven generalization frameworks (based on incremental and contrastive learning) cannot compensate for these feature-level shortcomings. To bridge this gap, we present KnowML, a framework that encodes attack domain knowledge directly into the feature space. For each attack family, our method employs LLMs to construct a corresponding Knowledge Graph (KG) from attack implementations. Symbolic reasoning is then applied over the KG to enumerate potential attack strategies and their compositions. The resulting Knowledge-Augmented Feature Space enables effective generalization even when trained exclusively on benign traffic, a capability beyond current approaches. Systematic empirical evaluations show that KnowML achieves up to 99% detection rates while maintaining false positive rates at or below 0.0137%, substantially outperforming contemporary feature-based baselines across diverse attack variants. Xin Fan Guo, Xinran Zheng, Albert Meroño-Peñuela, Lorenzo Cavallaro, Sergio Maffeis, Fabio Pierazzi |
EuroS&P | 2 |
| 2026 | Multi-Agent System for Orchestrating Retrieval and Multimodal Reasoning in Cloud Environments
Shuo Yang 0011, Xinran Zheng, Jinfeng Xu 0003, Jinze Li 0001, Edith C. H. Ngai |
INFOCOM | 2 |
| 2026 | CoLD: Collaborative Label Denoising Framework for Network Intrusion Detection
Shuo Yang 0011, Xinran Zheng, Jinze Li 0001, Jinfeng Xu 0003, Edith C. H. Ngai |
NDSS | 2 |
| 2026 | Robust federated intrusion detection under statistical heterogeneity
Xinran Zheng |
Comput. Networks | 1 |
| 2026 | TIF: Learning Temporal Invariance in Android Malware DetectorsabstractLearning-based Android malware detectors degrade over time due to natural distribution drift caused by malware variants and new families. This paper systematically investigates the challenges classifiers trained with empirical risk minimization (ERM) face against such distribution shifts and attributes their shortcomings to their inability to learnstablediscriminative features. Invariant learning theory offers a promising solution by encouraging models to generate stable representations across environments that expose the instability of the training set. However, the lack of prior environment labels, the diversity of drift factors, and low-quality representations caused by diverse families make this task challenging. To address these issues, we propose TIF, the first temporal invariant training framework for malware detection, which aims to enhance the ability of detectors to learn stable representations across time. TIF organizes environments based on application observation dates to reveal temporal drift, integrating specialized multi-proxy contrastive learning and invariant gradient alignment to generate and align environments with high-quality, stable representations. TIF can be seamlessly integrated into any learning-based detector. Experiments on a decade-long dataset show that TIF excels, particularly in early deployment stages, addressing real-world needs and outperforming state-of-the-art methods. Xinran Zheng, Shuo Yang 0011, Edith C. H. Ngai, Suman Jana, Lorenzo Cavallaro |
IEEE Trans. Software Eng. | 1 |
| 2025 | Mitigating Statistical Heterogeneity in Intrusion Detection Systems within Federated LearningabstractThe proliferation of the Internet of Things (IoT) is reshaping daily life while generating vast amounts of distributed user data, raising critical security and privacy concerns. Although federated learning-based intrusion detection systems (FL-based IDS) have shown promise in detecting potential attacks in such distributed environments, non-independent and identically distributed (non-i.i.d.) real-world applications’ data lead to considerable performance degradation. Current schemes mitigate this issue by sharing subsets of local clients’ data to guide consistency among clients; however, it inevitably compromises data privacy. In response to these challenges, we propose FedCKD, a federated framework that integrates well-designed Contrastive Knowledge Distillation for network intrusion detection. By utilizing a teacher-student model architecture on client devices, FedCKD aligns the feature maps of various models, ensuring consistent feature representation across clients without any lack of raw data. Furthermore, the incorporation of cluster contrastive learning enhances the separation between anomalous and normal instances, improving the model’s discriminative capabilities. We conduct extensive experiments on the NSL-KDD and UNSW-NB15 datasets, demonstrating that FedCKD outperforms existing methods, achieving over 90% f1-score (an increase of 3.67%). scenarios. Our approach effectively alleviates the impact of data heterogeneity while preserving data privacy, highlighting FedCKD as a robust and privacy-preserving solution for intrusion detection in IoT networks. Xinran Zheng, Shuo Yang 0011 |
GLOBECOM | 2 |
| 2025 | Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning AgentabstractMultimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the “hallucination” issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, we first construct Dyn-VQA dataset, consisting of three types of ``dynamic'' questions, which require complex knowledge retrieval strategies variable in query, tool, and time: (1) Questions with rapidly changing answers. (2) Questions requiring multi-modal knowledge. (3) Multi-hop questions. Experiments on Dyn-VQA reveal that existing heuristic mRAGs struggle to provide sufficient and precisely relevant knowledge for dynamic questions due to their rigid retrieval processes. Hence, we further propose the first self-adaptive planning agent for multimodal retrieval, **OmniSearch**. The underlying idea is to emulate the human behavior in question solution which dynamically decomposes complex multimodal questions into sub-question chains with retrieval action. Extensive experiments prove the effectiveness of our OmniSearch, also provide direction for advancing mRAG. Code and dataset will be open-sourced. Yangning Li, Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinran Zheng, Hui Wang 0030, Hai-Tao Zheng 0002, Fei Huang 0002, Jingren Zhou 0001, Philip S. Yu |
ICLR | 6 |
| 2025 | RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-CheckingabstractLarge Language Models (LLMs) hold significant potential for advancing fact-checking by leveraging their capabilities in reasoning, evidence retrieval, and explanation generation. However, existing benchmarks fail to comprehensively evaluate LLMs and Multimodal Large Language Models (MLLMs) in realistic misinformation scenarios. To bridge this gap, we introduce RealFactBench, a comprehensive benchmark designed to assess the fact-checking capabilities of LLMs and MLLMs across diverse real-world tasks, including Knowledge Validation, Rumor Detection, and Event Verification. RealFactBench consists of 6K high-quality claims drawn from authoritative sources, encompassing multimodal content and diverse domains. Our evaluation framework further introduces the Unknown Rate (UnR) metric, enabling a more nuanced assessment of models' ability to handle uncertainty and balance between over-conservatism and over-confidence. Extensive experiments on 7 representative LLMs and 4 MLLMs reveal their limitations in real-world fact-checking and offer valuable insights for further research. RealFactBench is publicly available at https://github.com/kalendsyang/RealFactBench.git. Shuo Yang 0011, Yuqin Dai, Xinran Zheng, Jinfeng Xu 0003, Jinze Li 0001, Zhenzhe Ying, Weiqiang Wang 0002, Edith C. H. Ngai |
ACM Multimedia | 4 |
| 2025 | Self-Supervised Adaptation Method to Concept Drift for Network Intrusion Detection
Shuo Yang 0011, Xinran Zheng, Jinze Li 0001, Jinfeng Xu 0003, Edith C. H. Ngai |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Multi-Scale Contrastive Attention Representation Learning for Encrypted Traffic ClassificationabstractEncrypted traffic classification is essential for network security and management. However, the encrypted nature makes it challenging to extract representative features from raw traffic data. Existing end-to-end methods ignore byte correlations within packets and potential correlations among packets, hindering the learning of real traffic semantics and leading to suboptimal performance. This paper proposes MsETC, a multi-scale contrastive attention representation learning method for encrypted traffic classification. MsETC divides the raw packet byte sequence into multi-scale patches and then extracts dual views for contrastive learning from both the inter-patch and intra-patch perspectives. This allows the model to capture correlations among bytes within a packet as well as the potential interactions between packets. Extensive experiments on real-world datasets demonstrate that the proposed method achieves superior classification performance with lower complexity. Shuo Yang 0011, Xinran Zheng, Jinze Li 0001, Jinfeng Xu 0003, Edith C. H. Ngai |
CIKM | 2 |
| 2024 | ReCDA: Concept Drift Adaptation with Representation Enhancement for Network Intrusion DetectionabstractThe deployment of learning-based models to detect malicious activities in network traffic flows is significantly challenged by concept drift. With evolving attack technology and dynamic attack behaviors, the underlying data distribution of recently arrived traffic flows deviates from historical empirical distributions over time. Existing approaches depend on a significant amount of labeled drifting samples to facilitate the deep model to handle concept drift, which faces labor-intensive manual labeling and the risk of label noise. In this paper, we propose ReCDA, a Concept Drift Adaptation method with Representation enhancement, which consists of a self-supervised representation enhancement stage and a weakly-supervised classifier tuning stage. Specifically, in the initial stage, ReCDA introduces drift-aware perturbation and representation alignment to facilitate the model in acquiring robust representations from drift-aware and drift-invariant perspectives. Moreover, in the subsequent stage, a meticulously crafted instructive sampling strategy and a robust representation constraint encourage the model to learn discriminative knowledge about benign and malicious activities during fine-tuning, thereby enhancing performance further. We conduct comprehensive evaluations on several benchmark datasets under varying degrees of concept drift. The experiment results demonstrate the superior adaptability and robustness of the proposed method. Shuo Yang 0011, Xinran Zheng, Jinze Li 0001, Jinfeng Xu 0003, Edith C. H. Ngai |
KDD | 2 |
| 2024 | An Efficient (t, n) Threshold Authentication Scheme for Vehicular Ad Hoc NetworksabstractVehicular Ad Hoc Network(VANET) is a hot topic, how to design an efficient authentication scheme is the core problem in this field. With the considerable progress in computing capability and side-channel attack capability, some seemingly safe third parties have also become less safe. Our paper focuses on weakening the use of trusted authority(TA) and improving the fault tolerance of authentication schemes. At first, all private keys are generated by users. Next, the ephemeral Diffie-Hellman key exchange scheme is used to construct a PID authentication. Moreover, we build a partial key delivery method without a secure channel. In the batch message authentication part, we use the Shamir secret share method to build a ($t,\ n$) threshold signature scheme that guarantees batch signatures cannot be tampered with fewer than$t$people involved. Finally, we make a comparison of time cost and communication cost with the other four schemes. Our scheme has a remarkable advantage over other schemes by comparing experimental data. Wenhui Kong, Pengfei Wen, Xinran Zheng, Shuo Yang 0011 |
WCNC | 3 |
| 2024 | STCA: Stacked Token-based Continuous Authentication Protocol for Zero Trust IoTabstractNetwork technology developments are blurring the security boundary of the Internet of Things (IoT), making it difficult for traditional border-based security architectures to cope with endless internal attacks. The Zero Trust Architecture (ZTA), with its core concept of “Never Trust, Always Verify”, effectively addresses this problem. Continuous Authentication (CA) is an indispensable component of the zero trust IoT. However, existing CA protocols are impractical to deploy on zero trust IoT due to their dependence on specific properties and thirst for resources. Therefore, this paper proposes a continuous authentication protocol, namely STCA, that uses stacked tokens to ensure the legitimacy of entities during a whole session. Through the theoretical analysis and simulation results, STCA resists several common attacks. The comparison analysis indicates that STCA has a better performance compared with related CA protocols, thereby demonstrating it is more suitable for zero trust IoT. Shuo Yang 0011, Xinran Zheng |
WCNC | 3 |
| 2024 | An Efficient and Scalable FHE-Based PDQ Scheme: Utilizing FFT to Design a Low Multiplication Depth Large-Integer Comparison AlgorithmabstractThe growing number of data privacy breaches and associated financial losses have driven the demand for private database queries. Clients typically submit queries that involve both search and computation operations, such as counting students under a certain age or calculating the BMI of employees above a specific age. Existing protocols often face limitations due to reliance on specific-purpose encryption schemes or multiple communication rounds between clients and servers. In this work, we present a unified framework utilizing fully homomorphic encryption techniques to efficiently and privately process queries with search and computation operations. Our contributions include a homomorphic encryption-based private comparison algorithm, called the layered comparison algorithm, which achieves a 2.6-6.6X performance improvement compared to algorithms from prior work; a fast Fourier transform-based preprocessing method enabling accurate large integer arithmetic operations in the encrypted domain; and a scalable database encoding method. Evaluation results demonstrate the practicality of our system, as it processes an aggregated query for a 1k-row encrypted database in approximately 4.53 seconds. Fahong Zhang 0002, Chen Yang 0005, Rui Zong, Xinran Zheng, Jianfei Wang 0003, Yishuo Meng |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | An Efficient and Adaptive Content Delivery System Based on Hybrid NetworkabstractNowadays, Content Delivery Network (CDN) is widely used for its convenience in providing services. However, the increasing demand for bandwidth puts tremendous pressure on CDN. Inspired by the great potential of periodic broadcasting to save bandwidth, we suggest an Efficient and Adaptive Content Delivery System (EACDS) to decrease traffic and shorten video content delivery delays. Specifically, we propose the Peak cutting and Valley filling for VBR (PVV) and the Exhaustive Periodic Broadcasting algorithm (EPB) to process videos and arrange slices with low bandwidth and delay. Furthermore, we introduce an enhanced version of EPB that supports Fast-Forwarding, namely FFB. To effectively deal with the constantly changing Internet, we also demonstrate an adaptive decision model that switches distribution schemes according to the online user scale. Extensive experiments show that the EACDS is excellent in many aspects. The PVV and decision model can save up to 57% of bandwidth and 54.7% of traffic respectively. Compared with some existing schemes, EPB allows the most negligible delay with the same bandwidth, and FFB supports fast-forwarding in periodic broadcasting. Linjie Nie, Shuo Yang 0011, Xinran Zheng |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Towards Efficient Blockchain-Based Cross-Domain Trust Reputation Management for LS-HetNetabstractThe large-scale heterogeneous network (LS-HetNet) provides seamless interconnection of everything by integrating different networks into a unified one, but it also comes with a significant risk of insider attacks that are difficult to address through traditional border-based security measures. Trust and reputation management (TRM) is a practical approach to solving this problem. However, past schemes are struggled to cope with the cross-domain architecture, scalability, and context dependency required by LS-HetNet, and have faced potential trust attacks. Hence, this paper proposes an efficient and robust cross-domain trust reputation model, MBC-Trust, to meet the above requirements simultaneously. MBC-Trust adopts a multi-level blockchain framework and a cloud-fog computing paradigm to support frequent and concurrent cross-domain trust sharing. Numerous network entities can access trust in an efficient and scalable manner through the collaboration of local and monitoring chains. In addition, a well-designed time decay function, recommendation trust factors, and other measures enhance the robustness of the model to attacks during trust evaluation and provide sensitive context awareness. Experiments demonstrate that MBC-Trust is highly efficient, robust, and scalable. Xinran Zheng, Shuo Yang 0011 |
GLOBECOM | 2 |
| 2023 | A Lightweight Approach for Network Intrusion Detection Based on Self-Knowledge DistillationabstractNetwork Intrusion Detection (NID) works as a kernel technology for the security network environment, obtaining extensive research and application. Despite enormous efforts by researchers, NID still faces challenges in deploying on resource-constrained devices. To improve detection accuracy while reducing computational costs and model storage simultaneously, we propose a lightweight intrusion detection approach based on self-knowledge distillation, namely LNet-SKD, which achieves the trade-off between accuracy and efficiency. Specifically, we carefully design the DeepMax block to extract compact representation efficiently and construct the LNet by stacking DeepMax blocks. Furthermore, considering compensating for performance degradation caused by the lightweight network, we adopt batchwise self-knowledge distillation to provide the regularization of training consistency. Experiments on benchmark datasets demonstrate the effectiveness of our proposed LNet-SKD, which outperforms existing state-of-the-art techniques with fewer parameters and lower computation loads. Shuo Yang 0011, Xinran Zheng, Zhengzhuo Xu |
ICC | 2 |
| 2023 | SF-IDS: An Imbalanced Semi-Supervised Learning Framework for Fine-Grained Intrusion DetectionabstractDeep learning-based fine-grained network intrusion detection systems (NIDS) enable different attacks to be responded to in a fast and targeted manner with the help of large-scale labels. However, the cost of labeling causes insufficient labeled samples. Also, the real fine-grained traffic shows a long-tailed distribution with great class imbalance. These two problems often appear simultaneously, posing serious challenges to fine-grained NIDS. In this work, we propose a novel semi-supervised fine-grained intrusion detection framework, SF-IDS, to achieve attack classification in the label-limited and highly class imbalanced case. We design a self-training backbone model called RI-1DCNN to boost the feature extraction by reconstructing the input samples into a multichannel image format. The uncertainty of the generated pseudo-labels is evaluated and used as a reference for pseudo-label filtering in combination with the prediction probability. To mitigate the effects of fine-grained class imbalance, we propose a hybrid loss function combining supervised contrastive loss and multi-weighted classification loss to obtain more compact intra-class features and clearer interclass intervals. Experiments show that the proposed SF-IDS achieves 3.01% and 2.71% Marco-F1 improvement on two classical datasets with 1% labeled, respectively. Xinran Zheng, Shuo Yang 0011 |
ICC | 1 |
| 2023 | A Reliable and Decentralized Trust Management Model for Fog Computing in Industrial IoTabstractFog computing facilitates low-latency computing and nearby storage of device data in the Industrial Internet of Things (IIoT). Frequent internal interactions between fog nodes can support complex industrial processes, creating a secure fog interaction environment imperative. However, decentralized, mobile, and heterogeneous fog nodes make static cryptography-based security schemes challenging to cope with internal attacks. Adding a trust management model (TMM) to the fog layer is an effective solution to address this issue. Existing fog-oriented TMMs mostly ignore the inconsistent behavior of nodes or blindly trust fake recommendations from highly trusted nodes, which makes TMMs vulnerable to trust attacks from individuals or groups. Also, diverse industrial scenarios require TMMs to sensitively capture changes in fog node behaviors and contexts, rather than relying solely on binary ratings. This paper proposes a reliable distributed trust management model, TFog, to secure industrial IoT fog node interactions. It provides all well-defined model components and their collaboration methods to operate independently on fog nodes in a scalable form. Novel recommendation filtering algorithms, bi-directional interaction rating, and multi-party trust update strategies are proposed to resist trust attacks. Experiments show that TFog has good sensitivity and convergence, also can effectively resist eight typical trust attacks. Xinran Zheng, Shuo Yang 0011 |
NOMS | 1 |
| 2023 | IBA: A secure and efficient device-to-device interaction-based authentication scheme for Internet of Things
Shuo Yang 0011, Xinran Zheng, Guining Liu |
Comput. Commun. | 2 |