VLDB 2026 Research / reviewers in the wild / expert
Ruizhi Xiao
dblp:294/2184
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9937-5774ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LuaReSym: Recovering Variable Liveness Ranges in Stripped Lua Bytecode via Multi-Stage Static AnalysisabstractLua is a lightweight scripting language widely adopted across diverse application domains. In practice, Lua applications are often distributed as compiled bytecode to protect intellectual property and improve loading efficiency. Existing Lua decompilers rely heavily on debug symbols embedded in bytecode to generate human-readable code. When debug symbols are stripped, these tools utilize heuristic-based methods to infer variable liveness ranges. However, existing heuristic methods often produce inaccurate predictions, reducing the readability of the decompilation results. Ruizhi Xiao, Jiakun Sun, Yuqing Shao, Shuyuan Jin |
ICPC | 2 |
| 2026 | CeeDet: A Class-Incremental Learning Method with Early-Exit Mechanism for Malicious Traffic Detection in IIoT
Jiakun Sun, Ruizhi Xiao, Shuyuan Jin |
PAKDD (1) | 3 |
| 2026 | InterpLog: Interpretable log-based anomaly detection assisting troubleshooting for system reliability
Ruizhi Xiao, Jiakun Sun, Shuyuan Jin |
J. Syst. Softw. | 1 |
| 2026 | Graph-Based Malicious Domain Name Detection: How to Use the Heuristic RelationsabstractDomain Name System is widely abused by various types of malicious campaigns. Recently, many graph learning models have been proposed to detect malicious domains based on the domain name resolution process and related data. These models focus on associations among domain names, which are defined as heuristic relations in this paper, and typically report an F1-score exceeding 0.95, indicating high detection accuracy achieved in real-world DNS applications. In order to explore how far we are from excellent graph-based malicious domain name detection methods, this paper conducts an in-depth analysis of six representative graph-based models on three experimental datasets and one real-world dataset. Our experiments focus on several aspects of model evaluation, including heuristic relation distributions, heuristic relation selection, feature extraction, graph reduction operation, and imbalance distribution of malicious domain names in the real world. The experimental results demonstrate that all these aspects have a significant impact on the detection performance, existing models are still relatively shallow in utilizing heuristic relations, and that all the studied models do not always work well as claimed. We further propose several possible future works that may contribute to achieving excellent performance in malicious domain name detection. Ruizhi Xiao, Jiakun Sun, Shuyuan Jin |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Cache Periscope: Gain insights into the global epidemic of malicious domains through DNS Cache
Jiakun Sun, Ruizhi Xiao, Shuyuan Jin |
Comput. Networks | 3 |
| 2024 | RCFG2Vec: Considering Long-Distance Dependency for Binary Code Similarity DetectionabstractBinary code similarity detection(BCSD), as a fundamental technique in software security, has various applications, including malware family detection, known vulnerability detection and code plagiarism detection. Recent deep learning-based BCSD approaches have demonstrated promising performance. However, they face two significant challenges that limit detection performance. First, most approaches that use sequence networks (like RNN and Transformer) utilize coarse-grained tokenization methods, which results in large vocabulary size and severe out-of-vocabulary (OOV) problem. Second, CFG-based methods typically use variants of graph convolutional networks, which only consider local structural information and discard long-distance dependencies between basic blocks. Jintian Lu, Ruizhi Xiao, Shuyuan Jin |
ASE | 3 |
| 2024 | MC-Det: Multi-channel representation fusion for malicious domain name detection
Ruizhi Xiao, Jiakun Sun, Shuyuan Jin |
Comput. Networks | 2 |
| 2024 | Secure and Real-Time Traceable Data Sharing in Cloud-Assisted IoTabstractCloud-assisted Internet of Things (IoT) has become an increasingly popular paradigm to greatly improve the performance of IoT applications by delegating the cloud to manage the massive IoT data. How to achieve secure and real-time traceable data sharing (STDS) is crucial in this paradigm, especially, a large amount of sensitive data produced by IoT devices needs to be stored or accessed to/from the clouds. This article proposes an STDS scheme, which leverages the acrlong DIFC model to allow data owners to not only securely and efficiently share their data produced by IoT devices with data users but also have the capability of tracking the data users’ identity with nonrepudiation based on the hash chain technique. Subsequently, the acrlong HLPN, acrlong SMT-Lib, and Z3 solver are used to formally analyze and verify STDS based on acrlong BMC technique to prove the correctness and security STDS. The formal analysis results show that STDS fulfills its intended security goals. Finally, the performance evaluation results have demonstrated the efficiency of STDS. Jintian Lu, Jiakun Sun, Ruizhi Xiao, Bolin Liao |
IEEE Internet Things J. | 4 |
| 2024 | ContexLog: Non-Parsing Log Anomaly Detection With All Information Preservation and Enhanced Contextual RepresentationabstractLogs are widely used in software to trace the runtime states and critical events. Log-based anomaly detection is crucial for software maintenance and reliability assurance. Existing log-based anomaly detection methods are suffering from imperfections of log parsing, the neglect of the log individual context, and the discarding of non-character tokens. In this paper, we propose ContexLog, a non-parsing log-based anomaly detection method with all information preservation and enhanced log contextual representation, to detect diverse anomalies effectively. Log messages are first grouped as sequences with different windowing techniques. To capture all log features, ContexLog tokenizes each log sequence and preserves all information, including character and non-character tokens. It then represents the log sequential context and individual context simultaneously to construct input for a Transformer encoder-based classification model. Experimental evaluations on real-world datasets and synthetic datasets demonstrate ContexLog outperforms existing methods in achieving accurate anomaly detection results, handling unseen logs to avoid log parsing imperfections, and utilizing non-character tokens to detect diverse anomalies. Ruizhi Xiao, Jintian Lu, Shuyuan Jin |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | AllInfoLog: Robust Diverse Anomalies Detection Based on All Log FeaturesabstractLarge-scale services are generating massive logs, which trace the runtime states and critical events. Anomaly detection via logs is critical for service maintenance and reliability assurance. Existing log-based anomaly detection methods make use of the limited information in log data, resulting in their incapability of detecting diverse anomalies related to unused log features. In this paper, we propose AllInfoLog, a robust log-based anomaly detection method taking advantage of all log information, to detect diverse types of anomalies. To capture all log features, AllInfoLog utilizes four encoders to extract semantic, parameter, time, and other feature embeddings, respectively. The embeddings of all log features are then combined to train an attention-based Bi-LSTM model to detect diverse anomalies. The experimental evaluations on real-world log datasets, synthetic datasets, and unstable log datasets demonstrate AllInfoLog outperforms the state-of-the-art log-based anomaly detection methods from aspects of performance and robustness, and has effectiveness to detect diverse types of anomalies. Ruizhi Xiao, Hao Chen 0133, Jintian Lu, Shuyuan Jin |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | MEMTD: Encrypted Malware Traffic Detection Using Multimodal Deep Learning
Jintian Lu, Jiakun Sun, Ruizhi Xiao, Shuyuan Jin |
ICWE | 4 |
| 2022 | DIFCS: A Secure Cloud Data Sharing Approach Based on Decentralized Information Flow Control
Jintian Lu, Jiakun Sun, Ruizhi Xiao, Shuyuan Jin |
Comput. Secur. | 3 |
| 2021 | Unsupervised Anomaly Detection Based on System LogsabstractThe anomaly detection based on rich and descriptive system logs is critical to securing information systems.Existing techniques rarely consider semantic information of logs in the detection, resulting in their incapability to handle unseen log events, neither further improve their detection rates.This paper proposes a CNN and LSTM based anomaly detection approach.It utilizes the meaning of log entries -the semantic information of logs in the detection, where the relations among short sequences are automatically learned.The results of comparative experiments demonstrate the effectiveness of the proposed approach on both stable(fixed format) and unstable(unseen, unfixed format) logs. Ruizhi Xiao, Shuyuan Jin |
SEKE | 2 |