VLDB 2026 Research / reviewers in the wild / expert
Hongyi Peng
dblp:36/2123
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-2720-2628ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRIORITI: scoring and categorization-based threat prioritization
Rajendra Patil 0001, Sivaanandh Muneeswaran, Vinay Sachidananda, Hongyi Peng, Gurusamy Mohan |
J. Supercomput. | 4 |
| 2024 | FedCal: Achieving Local and Global Calibration in Federated Learning via Aggregated Parameterized ScalerabstractFederated learning (FL) enables collaborative machine learning across distributed data owners, but data heterogeneity poses a challenge for model calibration. While prior work focused on improving accuracy for non-iid data, calibration remains under-explored. This study reveals existing FL aggregation approaches lead to sub-optimal calibration, and theoretical analysis shows despite constraining variance in clients’ label distributions, global calibration error is still asymptotically lower bounded. To address this, we propose a novel Federated Calibration (FedCal) approach, emphasizing both local and global calibration. It leverages client-specific scalers for local calibration to effectively correct output misalignment without sacrificing prediction accuracy. These scalers are then aggregated via weight averaging to generate a global scaler, minimizing the global calibration error. Extensive experiments demonstrate that FedCal significantly outperforms the best-performing baseline, reducing global calibration error by 47.66% on average. Hongyi Peng, Han Yu 0001, Xiaoli Tang 0001, Xiaoxiao Li 0001 |
ICML | 1 |
| 2024 | Efficient and Privacy-Preserving Feature Importance-Based Vertical Federated LearningabstractVertical Federated Learning (VFL) enables multiple data owners, each holding a different subset of features about a largely overlapping set of data samples, to collaboratively train a global model. The quality of data owners' local features affects the performance of the VFL model, which makes feature selection vitally important. However, existing feature selection methods for VFL either assume the availability of prior knowledge on the number of noisy features or prior knowledge on the post-training threshold of useful features to be selected, making them unsuitable for practical applications. To bridge this gap, we propose the Federated Stochastic Dual-Gate based Feature Selection (FedSDG-FS) approach. It consists of a Gaussian stochastic dual-gate to efficiently approximate the probability of a feature being selected. FedSDG-FS further designs a local embedding perturbation approach to achieve differential privacy for local training data. To reduce overhead, we propose a feature importance initialization method based on Gini impurity, which can accomplish its goals with only two parameter transmissions between the server and the clients. The enhanced version, FedSDG-FS++, protects the privacy for both the clients' training data and the server's labels through Partially Homomorphic Encryption (PHE) without relying on a trusted third-party. Theoretically, we analyze the convergence rate, privacy guarantees and security analysis of our methods. Extensive experiments on both synthetic and real-world datasets show that FedSDG-FS and FedSDG-FS++ significantly outperform existing approaches in terms of achieving more accurate selection of high-quality features as well as improving VFL performance in a privacy-preserving manner. Anran Li 0001, Ju Jia, Hongyi Peng, Lan Zhang 0002, Anh Tuan Luu, Han Yu 0001, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Privacy-Preserving Data Selection for Horizontal and Vertical Federated LearningabstractFederated learning (FL) enables distributed participants to collaboratively train a machine learning model without accessing to their local data. In FL systems, the selection of training samples has a significant impact on model performances, e.g., selecting participants whose datasets have low-quality samples, features would result in low accuracy, unstable models. In this work, we aim to solve the problem that selects a collection of high-quality training samples for a given FL task under a monetary budget. We propose a holistic design to efficiently select high-quality samples while preserve the privacy of participants’ local data, the server’s label set. We propose an efficient hierarchical sample selection mechanism to select relevant clients, their samples before training for horizontal federated learning (HFL). It uses the determinantal point process (DPP) to select both the statistical homogenous, content diverse clients, samples. Besides, we propose a private set intersection (PSI) based scheme to filter relevant features for the target VFL task. Finally, during training, an erroneous-aware importance based selection is proposed to dynamically select important clients, samples to accelerate model convergence. We verify the merits of our proposed solution with extensive experiments on a real AIoT system with 50 clients. The experimental results validate that our solution achieves accurate, efficient selection of high-quality data, consequently an FL model with a faster convergence speed, higher accuracy. Lan Zhang 0002, Anran Li 0001, Hongyi Peng, Xiang-Yang Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | FedSDG-FS: Efficient and Secure Feature Selection for Vertical Federated LearningabstractVertical Federated Learning (VFL) enables multiple data owners, each holding a different subset of features about largely overlapping sets of data sample(s), to jointly train a useful global model. Feature selection (FS) is important to VFL. It is still an open research problem as existing FS works designed for VFL either assumes prior knowledge on the number of noisy features or prior knowledge on the post-training threshold of useful features to be selected, making them unsuitable for practical applications. To bridge this gap, we propose the Federated Stochastic Dual-Gate based Feature Selection (FedSDG-FS) approach. It consists of a Gaussian stochastic dual-gate to efficiently approximate the probability of a feature being selected, with privacy protection through Partially Homomorphic Encryption without a trusted third-party. To reduce overhead, we propose a feature importance initialization method based on Gini impurity, which can accomplish its goals with only two parameter transmissions between the server and the clients. Extensive experiments on both synthetic and real-world datasets show that FedSDG-FS significantly outperforms existing approaches in terms of achieving accurate selection of high-quality features as well as building global models with improved performance. Anran Li 0001, Hongyi Peng, Lan Zhang 0002, Qing Guo 0005, Han Yu 0001, Yang Liu 0003 |
INFOCOM | 2 |
| 2023 | ThreatLand: Extracting Intelligence from Audit Logs via NLP methodsabstractThreat intelligence and hunting using various logs has evolved into a crucial component of remaining aware of the ever-changing threat landscape. Given the critical need to extract useful intelligence from logs, existing techniques either focus exclusively on isolated records, ignoring correlation and the overall threat scenario, or require significant effort to filter and correlate threat records. Additionally, searching for and matching threat behaviors in logs often involves non-trivial human query construction, impeding fast threat hunting. To address this gap, we present ThreatLand, a system that extracts highlevel intelligence and structured threat patterns from audit logs automatically. ThreatLand is composed of three components (1) A lightweight and accurate NLP pipeline that extracts structured meta-data from alert descriptions and generates a heterogeneous graph that depicts the entire threat scenario. (2) A query execution engine that is both fast and efficient, based on a graphical database. (3) A graphical user interface (GUI) that offers various sorts of interactivity to aid intelligence exploration.We have evaluated the ThreatLand over the dataset containing 9240 real-time EDR alerts collected for the threat events over an enterprise setup in the lab. As a result, ThreatLand presents high-level insights from the alert logs and extracts the valuable threat patterns. Vinay Sachidananda, Rajendra Patil 0001, Hongyi Peng, Yang Liu 0003, Kwok-Yan Lam |
PST | 3 |
| 2023 | FedCSS: Joint Client-and-Sample Selection for Hard Sample-Aware Noise-Robust Federated LearningabstractFederated Learning (FL) enables a large number of data owners (a.k.a. FL clients) to jointly train a machine learning model without disclosing private local data. The importance of local data samples to the FL model vary widely. This is exacerbated by the presence of noisy data, which exhibit large losses similar to important (hard) samples. Currently, there lacks an FL approach that can effectively distinguish hard samples (which are beneficial) from noisy samples (which are harmful). To bridge this gap, we propose the Federated Client and Sample Selection (FedCSS) approach. It is a bilevel optimization approach for FL client-and-sample selection to achieve hard sample-aware noise-robust learning in a privacy preserving manner. It performs meta-learning based online approximation to iteratively update global FL models, select the most positively influential samples and deal with training data noise. Theoretical analysis shows that it is guaranteed to converge in an efficient manner. Experimental comparison against six state-of-the-art baselines on five real-world datasets in the presence of data noise and heterogeneity shows that it achieves up to 26.4% higher test accuracy, while saving communication and computation costs by at least 41.5% and 1.2%, respectively. Anran Li 0001, Jiabao Guo, Hongyi Peng, Qing Guo 0005, Han Yu 0001 |
Proc. ACM Manag. Data | 4 |
| 2022 | Peekaboo: Hide and Seek with Malware Through Lightweight Multi-feature Based Lenient Hybrid Approach
Mingchang Liu, Vinay Sachidananda, Hongyi Peng, Rajendra Patil 0001, Sivaanandh Muneeswaran, Gurusamy Mohan |
ICICS | 3 |
| 2022 | ODDITY: An Ensemble Framework Leverages Contrastive Representation Learning for Superior Anomaly Detection
Hongyi Peng, Vinay Sachidananda, Teng Joon Lim, Rajendra Patil 0001, Mingchang Liu, Sivaanandh Muneeswaran, Gurusamy Mohan |
ICICS | 1 |
| 2022 | LOG-OFF: A Novel Behavior Based Authentication Compromise Detection ApproachabstractPassword-based authentication system has been praised for its user-friendly, cost-effective, and easily deployable features. It is arguably the most commonly used security mechanism for various resources, services, and applications. On the other hand, it has well-known security flaws, including vulnerability to guessing attacks. Present state-of-the-art approaches have high overheads, as well as difficulties and unreliability during training, resulting in a poor user experience and a high false positive rate. As a result, a lightweight authentication compromise detection model that can make accurate detection with a low false positive rate is required.In this paper we propose – LOG-OFF – a behavior-based authentication compromise detection model. LOG-OFF is a lightweight model that can be deployed efficiently in practice because it does not include a labeled dataset. Based on the assumption that the behavioral pattern of a specific user does not suddenly change, we study the real-world authentication traffic data. The dataset contains more than 4 million records. We use two features to model the user behaviors, i.e., consecutive failures and login time, and develop a novel approach. LOG-OFF learns from the historical user behaviors to construct user profiles and makes probabilistic predictions of future login attempts for authentication compromise detection. LOG-OFF has a low false positive rate and latency, making it suitable for real-world deployment. In addition, it can also evolve with time and make more accurate detection as more data is being collected. Mingchang Liu, Vinay Sachidananda, Hongyi Peng, Rajendra Patil 0001, Sivaanandh Muneeswaran, Gurusamy Mohan |
PST | 3 |
| 2022 | Hiatus: Unsupervised Generative Approach for Detection of DoS and DDoS Attacks
Sivaanandh Muneeswaran, Vinay Sachidananda, Rajendra Patil 0001, Hongyi Peng, Mingchang Liu, Gurusamy Mohan |
SecureComm | 4 |
| 2022 | MARK: Fill in the blanks through a JointGAN based data augmentation for network anomaly detection
Rajendra Patil 0001, Vinay Sachidananda, Hongyi Peng, Akshay Sachdeva, Gurusamy Mohan |
Comput. Secur. | 3 |
| 2020 | A novel approach for detecting vulnerable IoT devices connected behind a home NATabstractTelecommunication service providers (telcos) are exposed to cyber-attacks executed by compromised IoT devices connected to their customers’ networks. Such attacks might have severe effects on the attack target, as well as the telcos themselves. To mitigate those risks, we propose a machine learning-based method that can detect specific vulnerable IoT device models connected behind a domestic NAT, thereby identifying home networks that pose a risk to the telcos infrastructure and service availability. To evaluate our method, we collected a large quantity of network traffic data from various commercial IoT devices in our lab and compared several classification algorithms. We found that (a) the LGBM algorithm produces excellent detection results, and (b) our flow-based method is robust and can handle situations for which existing methods used to identify devices behind a NAT are unable to fully address, e.g., encrypted, non-TCP or non-DNS traffic. To promote future research in this domain we share our novel labeled benchmark dataset. Yair Meidan, Vinay Sachidananda, Hongyi Peng, Racheli Sagron, Yuval Elovici, Asaf Shabtai |
Comput. Secur. | 3 |
| 2013 | Optimal gene subset selection using the modified SFFS algorithm for tumor classification
Hongyi Peng, Yinlian Fu, Jinshan Liu, Chunfu Jiang |
Neural Comput. Appl. | 1 |
| 2007 | Handling of incomplete data sets using ICA and SOM in data mining
Hongyi Peng, Siming Zhu |
Neural Comput. Appl. | 1 |