VLDB 2026 Research / reviewers in the wild / expert
Baihe Ma
dblp:238/5753
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-4167-2797ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Joint Trajectory Obfuscation and Pseudonym Swapping Mechanism Avoiding Extra Privacy Cost
Baihe Ma, Xu Wang 0004, Guangsheng Yu, Yanna Jiang, Suirui Zhu, Bo Liu 0001, Ying He 0011, Wei Ni 0001, Ren Ping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | Beyond Spatial Privacy: Protecting Trajectories With Spatio-Temporal Differential PrivacyabstractSpatio-temporal trajectories carry identifying information and are vulnerable to privacy breaches. Existing studies predominantly focus on the spatial domain. The temporal aspect remains underexplored, leaving privacy risks unaddressed. This paper highlights these risks by introducing a new trajectory matching model, ST-ATT, which leverages attention-enhanced Long Short-Term Memory (LSTM) to effectively capture the spatio-temporal correlations within trajectories. ST-ATT excels in identifying similar trajectories. To defend against linkage attacks on spatio-temporal trajectories, including advanced models like ST-ATT, we propose a novel Differential Privacy (DP) mechanism specifically designed to address the privacy risks. We reveal that the privacy budget and violation probability for each spatial point explicitly depend on earlier timestamps. The privacy budget can be flexibly redistributed between spatial and temporal domains without compromising overall privacy. This mechanism complies with DP, even when spatio-temporal points are reordered due to perturbation. Experiments show that ST-ATT can accurately identify spatio-temporal trajectories perturbed by the existing DP methods adding noise solely to the spatial domain. The proposed spatio-temporal DP mechanism resists ST-ATT, highlighting the need for considering spatio-temporal correlations to ensure robust privacy protection in spatio-temporal trajectories. Suirui Zhu, Xin Yuan 0004, Baihe Ma, Wei Ni 0001, Wenjie Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Split UnlearningabstractWe introduce Split Unlearning, a novel machine unlearning technology designed for Split Learning (SL), enabling the first-ever implementation of Sharded, Isolated, Sliced, and Aggregated (SISA) unlearning in SL frameworks. Particularly, the tight coupling between clients and the server in existing SL frameworks results in frequent bidirectional data flows and iterative training across all clients, violating the ''Isolated'' principle and making them struggle to implement SISA for independent and efficient unlearning. To address this, we propose SplitWiper with a new one-way-one-off propagation scheme, which leverages the inherently ''Sharded'' structure of SL and decouples neural signal propagation between clients and the server, enabling effective SISA unlearning even in scenarios with absent clients. We further design SplitWiper+ to enhance client label privacy, which integrates differential privacy and label expansion strategy to defend the privacy of client labels against the server and other potential adversaries. Experiments across diverse data distributions and tasks demonstrate that SplitWiper achieves 0% accuracy for unlearned labels, and 8% better accuracy for retained labels than non-SISA unlearning in SL. Moreover, the one-way-one-off propagation maintains constant overhead, reducing computational and communication costs by 99%. SplitWiper+ preserves 90% of label privacy when sharing masked labels with the server. Yanna Jiang, Guangsheng Yu, Qin Wang 0008, Xu Wang 0004, Baihe Ma, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001 |
CCS | 5 |
| 2025 | Exploiting attribute correlation for reconstruction attacks on differentially private multi-attributed data
Yanna Jiang, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001 |
J. Inf. Secur. Appl. | 2 |
| 2025 | CAN-Trace Attack: Exploit CAN Messages to Uncover Driving TrajectoriesabstractDriving trajectory data remains vulnerable to privacy breaches despite existing mitigation measures. Traditional methods for detecting driving trajectories typically rely on map-matching the path using Global Positioning System (GPS) data, which is susceptible to GPS data outage. This paper introduces CAN-Trace, a novel privacy attack mechanism that leverages Controller Area Network (CAN) messages to uncover driving trajectories, posing a significant risk to drivers’ long-term privacy. A new trajectory reconstruction algorithm is proposed to transform the CAN messages, specifically vehicle speed and accelerator pedal position, into weighted graphs accommodating various driving statuses. CAN-Trace identifies driving trajectories using graph-matching algorithms applied to the created graphs in comparison to road networks. We also design a new metric to evaluate matched candidates, which allows for potential data gaps and matching inaccuracies. Empirical validation under various real-world conditions, encompassing different vehicles and driving regions, demonstrates the efficacy of CAN-Trace: it achieves an attack success rate of up to 90.59% in the urban region, and 99.41% in the suburban region. Xiaojie Lin, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Ying He 0011, Wei Ni 0001, Ren Ping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | DPAC: A New Data-Centric Privacy-Preserving Access Control Model
Xu Wang 0004, Baihe Ma, Ren Ping Liu 0001, Ian J. Oppermann |
ProvSec (2) | 2 |
| 2024 | Preventing harm to the rare in combating the malicious: A filtering-and-voting framework with adaptive aggregation in federated learningabstractThe distributed nature of Federated Learning (FL) introduces security vulnerabilities and issues related to the heterogeneous distribution of data. Traditional FL aggregation algorithms often mitigate security risks by excluding outliers, which compromises the diversity of shared information. In this paper, we introduce a novel filtering-and-voting framework that adeptly navigates the challenges posed by non-iid training data and malicious attacks on FL. The proposed framework integrates a filtering layer for defensive measures against the intrusion of malicious models and a voting layer to harness valuable contributions from diverse participants. Moreover, by employing Deep Reinforcement Learning (DRL) for dynamic aggregation weight adjustment, we ensure the optimized aggregation of participant data, enhancing the diversity of information used for aggregation and improving the performance of the global model. Experimental results demonstrate that the proposed framework presents superior accuracy over traditional and contemporary FL aggregation methods as diverse models are utilized. It also shows robust resistance against malicious poisoning attacks. Yanna Jiang, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Caijun Sun, Wei Ni 0001, Ren Ping Liu 0001 |
Neurocomputing | 2 |
| 2024 | ByCAN: Reverse Engineering Controller Area Network (CAN) Messages From Bit to Byte LevelabstractAs the primary standard protocol for modern cars, the controller area network (CAN) is a critical research target for automotive cybersecurity threats and autonomous applications. As the decoding specification of CAN is a proprietary black-box maintained by original equipment manufacturers (OEMs), conducting related research and industry developments can be challenging without a comprehensive understanding of the meaning of CAN messages. In this article, we propose a fully automated reverse-engineering system, named ByCAN, to reverse engineer CAN messages. ByCAN outperforms the existing research by introducing byte-level clusters and integrating multiple features at both the byte and bit levels. ByCAN employs the clustering and template matching algorithms to automatically decode the specifications of CAN frames without the need for prior knowledge. Experimental results demonstrate that ByCAN achieves high accuracy in slicing and labeling performance, i.e., the identification of CAN signal boundaries and labels. In the experiments, ByCAN achieves slicing accuracy of 80.21%, slicing coverage of 95.21%, and labeling accuracy of 68.72% for the general labels when analysing the real-world CAN frames. Xiaojie Lin, Baihe Ma, Xu Wang 0004, Guangsheng Yu, Ying He 0011, Ren Ping Liu 0001, Wei Ni 0001 |
IEEE Internet Things J. | 2 |
| 2023 | MT-CNN: A Classification Method of Encrypted Traffic Based on Semi-Supervised LearningabstractDeep learning methods have become the preferred solution for encrypted traffic classification. However, the application of neural networks in encrypted traffic classification has encountered the following limitations: 1) Deep learning models have dependencies on large-scale and well-labeled datasets. 2) most deep learning models have high hardware requirements and require a large amount of CPU and GPU for computation. These limitations seriously hinder the development of encrypted traffic research. In this paper, we propose a new lightweight semi-supervised learning classifier to solve these problems. To reduce the dependence of the model on CPU and GPU, we have designed a lightweight encrypted traffic classifier based on CNN(Convolutional Neural Networks). It can run on raspberry pi with low hardware requirements. Then we combine the classifier with the Mean Teacher framework, which we call MT-CNN. By using the semi-supervised learning framework, we successfully reduced the number of labeled samples during model training. To fully preserve traffic information, we convert traffic data into grayscale images as input. We used a small-scale dataset for experiments on raspberry pi. The experimental results showed that the accuracy of MT-CNN still reached 96.83% even when only 5% of the labeled data was used. Kaichao Shi, Yong Zeng 0002, Baihe Ma, Jianfeng Ma 0001 |
GLOBECOM | 3 |
| 2022 | Multi-layer Reverse Engineering System for Vehicular Controller Area Network MessagesabstractThe undisclosed Controller Area Network (CAN) decoding specification is important to the in-vehicle network (IVN) research for both industry and academia. Researchers have developed several CAN reverse engineering systems to predict signal boundaries and labels in order to map out CAN signal decoding specifications. Existing works mainly use one parameter (i.e., bit flip rate) to determine CAN signals boundary, which results in biased slicing and labelling of CAN signals. In this paper, we propose a multi-layer CAN reverse engineering system to cluster signal boundary at byte-level and label sliced CAN signal blocks at bit-level. The proposed system avoids biased signal slicing and labelling by introducing multiple parameters in signal classification, while existing works only use the bit flip rate and the number of unique value. The feasibility and adaptability of the proposed system is assessed by deploying it into a web application as a functionality module. We evaluate the proposed system with CAN messages from real cars. Compared with existing reverse engineering models, the proposed system introduces multi-layer signal processing to avoid over-slicing and over-labelling problem. Xiaojie Lin, Baihe Ma, Xu Wang 0004, Ying He 0011, Ren Ping Liu 0001, Wei Ni 0001 |
CSCWD | 2 |
| 2022 | Trajectory Obfuscation and Detection in Internet-of-vehiclesabstractIn Internet-of-vehicles, vehicles cooperate with each other by transmitting Internet-of-vehicles and location-based service (LBS) providers optimize services by analyzing trajectory data collected from drivers. Nevertheless, illegal trajectories generated by attackers or malicious drivers can obfuscate the process of analysis and breach the quality of service. Some mechanisms protect drivers’ location privacy by using obfuscation-based schemes. Obfuscation-based mechanisms report LBS with obfuscated trajectories data rather than actual trajectories, which increases difficulties to detect illegal trajectories accurately. This paper focuses on detecting illegal trajectories when all drivers employ obfuscation-based mechanisms to protect location privacy. In this paper, we propose a dynamic obfuscation mechanism in road networks based on Geo-indistinguishability to dynamically protect drivers’ location privacy. Considering personalization in road networks, we also propose a classification mechanism to detect illegal trajectories in road networks. Illegal trajectories are generated based on real trajectories to simulate actions of malicious drivers and attackers. Experiment results in real road networks show that the classifier can detect illegal obfuscated trajectories with at least 94% Area Under the Curve (AUC) score, which outperforms than existing works in road networks. Yueyao Zhao, Baihe Ma, Ziwen Wang 0005, Yong Zeng 0002, Jianfeng Ma 0001 |
CSCWD | 2 |
| 2022 | Leveraging Byte-Level Features for LSTM-based Anomaly Detection in Controller Area NetworksabstractThe legacy design of the Controller Area Network (CAN) weakens the encryption and authentication of the In-Vehicle Networks (IVN). Anomaly detection systems, e.g. the Long-Short Term Memory (LSTM) based Intrusion Detection System (IDS), are employed to remedy the defection of CAN. Existing works feed the LSTM-based IDS with the byte values of the data payload of CAN to train and test the LSTM model. In this paper, we propose an LSTM-based IDS leveraging byte-level features, i.e., byte flip rate, byte-level change rage, and byte-level distinct value rate, to augment the sensitivity of proposed LSTM-based IDS when distinguishing malicious CAN messages. By using the byte-level signal features, the proposed system achieves high accuracy with a small size of the training dataset. The experiment results show that the model with the byte-level features can achieve a performance gain of the$F$1Score up to 20% over the model without the byte-level features. Lixue Liang, Xiaojie Lin, Baihe Ma, Xu Wang 0004, Ying He 0011, Ren Ping Liu 0001, Wei Ni 0001 |
GLOBECOM | 3 |
| 2022 | New Cloaking Region Obfuscation for Road Network-Indistinguishability and Location PrivacyabstractThe development of location-based services (LBS) leads to the rapid growth of location data, potentially increasing the threat to location privacy. Existing location obfuscation techniques focus on two-dimensional (2D) planar areas and overlook the features of road networks. In this paper, we leverage differential privacy and propose a new notion of Road Network-Indistinguishability (RN-Indistinguishability) to measure the indistinguishability of locations in road networks. With the RN-Indistinguishability, we design a Cloaking Region Obfuscation (CRO) mechanism to protect the location privacy of vehicles on roads. With the CRO mechanism, vehicle locations in a cloaking region are obfuscated following the same obfuscation distribution. The proposed CRO mechanism is proved to achieve RN-Indistinguishability and can be generalized with road network features holding the triangle inequality. Comprehensive experiments show that the CRO mechanism outperforms existing 2D obfuscation mechanisms in real-world road networks. Baihe Ma, Xiaojie Lin, Xu Wang 0004, Bin Liu 0028, Ying He 0011, Wei Ni 0001, Ren Ping Liu 0001 |
RAID | 1 |
| 2022 | Personalized Location Privacy With Road Network-IndistinguishabilityabstractThe proliferation of location-based services (LBS) leads to increasing concern about location privacy. Location obfuscation is a promising privacy-preserving technique but yet to be adequately tailored for vehicles in road networks. Existing obfuscation schemes are based primarily on the Euclidean distances and can lead to infeasible results, e.g., off-road locations. In this paper, we define Road Network-Indistinguishability (RN-I) to evaluate obfuscation-based location privacy-preserving schemes in road networks. To protect drivers’ location privacy in road networks, we propose a Personalized Location Privacy-Preserving (PLPP) scheme and prove it achieves RN-I. The PLPP scheme employs a dual-obfuscation algorithm, consisting of a connection perturbation and an interval perturbation, to obfuscate on-road locations. An efficient personalization algorithm is designed for the PLPP scheme to fine-tune location privacy budgets for capturing drivers’ sensitive locations and privacy requirements. Experiments upon two real-world datasets confirm the location privacy-preserving capability, data utility, and efficiency of the proposed PLPP scheme. Baihe Ma, Xu Wang 0004, Wei Ni 0001, Ren Ping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |