VLDB 2026 Research / reviewers in the wild / expert
Zhenzhou Tian
dblp:145/4398
· DBLP profile ↗
23ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0001-7608-8908ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Security and privacy · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust vulnerability detection with limited data via training-efficient adversarial reprogramming
Zhenzhou Tian, Yunpeng Hui, Jiaze Sun, Yanping Chen 0006, Lingwei Chen |
Autom. Softw. Eng. | 1 |
| 2026 | When fixes teach: Repair-aware contrastive learning for optimization-resilient binary vulnerability detection
Zhenzhou Tian, Ming Fan 0002, Jiaze Sun, Yanping Chen 0006, Lingwei Chen |
J. Syst. Archit. | 1 |
| 2025 | SolBERT: Advancing solidity smart contract similarity analysis via self-supervised pre-training and contrastive fine-tuning
Zhenzhou Tian, Yudong Teng, Xianqun Ke, Yanping Chen 0006, Lingwei Chen |
Inf. Softw. Technol. | 1 |
| 2025 | HardVD: High-capacity cross-modal adversarial reprogramming for data-efficient vulnerability detection
Zhenzhou Tian, Haojiang Li, Hanlin Sun, Yanping Chen 0006, Lingwei Chen |
Inf. Sci. | 1 |
| 2025 | Towards cost-efficient vulnerability detection with cross-modal adversarial reprogramming
Zhenzhou Tian, Yudong Teng, Jiaze Sun, Yanping Chen 0006, Lingwei Chen |
J. Syst. Softw. | 1 |
| 2024 | Enhancing vulnerability detection via AST decomposition and neural sub-tree encoding
Zhenzhou Tian, Binhui Tian, Jiajun Lv, Yanping Chen 0006, Lingwei Chen |
Expert Syst. Appl. | 1 |
| 2024 | Function-Level Code Obfuscation Detection Through Self-Attention-Guided Multi-Representation FusionabstractMalware developers often employ code obfuscation techniques to conceal their malicious functionality, making it challenging to detect and analyze such software. While various de-obfuscation techniques exist, the majority of them require prior knowledge of the obfuscation tools and techniques in use. Identifying the specific obfuscation tools or algorithms applied to the obfuscated code is thus of vital importance, which, however, typically demands in-depth expert knowledge and substantial efforts. Therefore, this paper presents DeObA, a deep learning (DL) driven approach for the precise and efficient detection of obfuscation algorithms on the fine-grained function-level code snippets. To comprehensively capture unique patterns or features of different obfuscation algorithms from code, DeObA works on multiple distinct code views, encompassing token sequences, abstract syntax trees (AST) and program dependency graphs (PDG), which will reflect the code’s lexical morphology, syntactic and structural aspects. After individually collecting obfuscation-indicative features with well-matched DL encoder from each code view, a self-attention-based fusion strategy is performed on these features to produce an integrated, dense, yet feature-rich vector. This vector is then fed into a softmax classification layer for prediction. Due to the lack of a moderately sized dataset, a large obfuscation corpus is curated with 7 different obfuscation tools and a total of 12 obfuscation algorithms on 39,070 C/C[Formula: see text] functions. The experimental evaluations conducted on the dataset exhibit a distinguished detection performance of DeObA, which achieve accuracy rates of 99.90% and 99.19% on the obfuscation tool detection and obfuscation algorithm detection tasks, respectively. The ablation study also confirms the active role of considering multiple distinct code views and the effectiveness of the designed self-attention-based fusion strategy. Zhenzhou Tian, Ruikang He, Hongliang Zhao, Lingwei Chen |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Differential testing solidity compiler through deep contract manipulation and mutation
Zhenzhou Tian, Fanfan Wang, Yanping Chen 0006, Lingwei Chen |
Softw. Qual. J. | 1 |
| 2022 | Ethereum Smart Contract Representation Learning for Robust Bytecode-Level Similarity DetectionabstractSmart contracts are programs that run on a blockchain, where Ethereum is one of the most popular ones supporting them.Due to the fact that they are immutable, it is essential to design smart contracts bug-free before they are deployed.However, various defects have been found in the deployed smart contracts, causing huge economic losses and lowing people's trust.Writing secure smart contracts is far from trivial, where developers tend to engage in reliable resources or social coding platforms to reuse code.This leads to a large number of similar contracts with potential security risks.Therefore, detecting similarity of smart contracts helps to avoid vulnerabilities, identify threats, and improve the security of Ethereum.In this paper, we design a learning-effective and costefficient model, called SmartSD, for Ethereum smart contract similarity detection.Different from the current research efforts, SmartSD is performed on a bytecode level and leverages deep neural networks to learn the latent representations from the opcode sequences for smart contract bytecodes, where the representation learning and similarity measurement are supervised via siamese neural networks.The experimental evaluations demonstrate that SmartSD outperforms EClone's 93.27% accuracy, achieving 98.37% high detection accuracy and 0.9850 F1-score, which is computationally tractable and effectively mitigates the interference caused by compilers. Zhenzhou Tian, Zhongmin Wang 0001, Yanping Chen 0006, Lingwei Chen |
SEKE | 1 |
| 2022 | Landscape estimation of solidity version usage on Ethereum via version identification
Zhenzhou Tian, Zhongmin Wang 0001, Yanping Chen 0006, Hong Xia, Lingwei Chen |
Int. J. Intell. Syst. | 1 |
| 2021 | Fine-Grained Obfuscation Scheme Recognition on Binary Code
Zhenzhou Tian, Hengchao Mao, Jinrui Li |
ICDF2C | 1 |
| 2020 | Neural Representation Learning Based Binary Code Authorship Attribution
Zhongmin Wang 0001, Zhenzhou Tian |
ICDF2C | 3 |
| 2020 | Plagiarism Detection of Multi-threaded Programs using Frequent Behavioral Pattern Mining
Qing Wang 0030, Zhenzhou Tian, Cong Gao 0002, Lingwei Chen |
SEKE | 2 |
| 2020 | Towards Fine-Grained Compiler Identification with Neural Modeling
Borun Xie, Zhenzhou Tian, Cong Gao 0002, Lingwei Chen |
SEKE | 2 |
| 2020 | Plagiarism Detection of Multi-threaded Programs Using Frequent Behavioral Pattern MiningabstractSoftware dynamic birthmark techniques construct birthmarks using the captured execution traces from running the programs, which serve as one of the most promising methods for obfuscation-resilient software plagiarism detection. However, due to the perturbation caused by non-deterministic thread scheduling in multi-threaded programs, such dynamic approaches optimized for sequential programs may suffer from the randomness in multi-threaded program plagiarism detection. In this paper, we propose a new dynamic thread-aware birthmark FPBirth to facilitate multi-threaded program plagiarism detection. We first explore dynamic monitoring to capture multiple execution traces with respect to system calls for each multi-threaded program under a specified input, and then leverage the Apriori algorithm to mine frequent patterns to formulate our dynamic birthmark, which can not only depict the program’s behavioral semantics, but also resist the changes and perturbations over execution traces caused by the thread scheduling in multi-threaded programs. Using FPBirth, we design a multi-threaded program plagiarism detection system. The experimental results based on a public software plagiarism sample set demonstrate that the developed system integrating our proposed birthmark FPBirth copes better with multi-threaded plagiarism detection than alternative approaches. Compared against the dynamic birthmark System Call Short Sequence Birthmark (SCSSB), FPBirth achieves 12.4%, 4.1% and 7.9% performance improvements with respect to union of resilience and credibility (URC), F-Measure and matthews correlation coefficient (MCC) metric, respectively. Zhenzhou Tian, Qing Wang 0030, Cong Gao 0002, Lingwei Chen, Dinghao Wu |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2018 | Android Malware Familial Classification and Representative Sample Selection via Frequent Subgraph AnalysisabstractThe rapid increase in the number of Android malware poses great challenges to anti-malware systems, because the sheer number of malware samples overwhelms malware analysis systems. The classification of malware samples into families, such that the common features shared by malware samples in the same family can be exploited in malware detection and inspection, is a promising approach for accelerating malware analysis. Furthermore, the selection of representative malware samples in each family can drastically decrease the number of malware to be analyzed. However, the existing classification solutions are limited because of the following reasons. First, the legitimate part of the malware may misguide the classification algorithms because the majority of Android malware are constructed by inserting malicious components into popular apps. Second, the polymorphic variants of Android malware can evade detection by employing transformation attacks. In this paper, we propose a novel approach that constructs frequent subgraphs (fregraphs) to represent the common behaviors of malware samples that belong to the same family. Moreover, we propose and develop FalDroid, a novel system that automatically classifies Android malware and selects representative malware samples in accordance with fregraphs. We apply it to 8407 malware samples from 36 families. Experimental results show that FalDroid can correctly classify 94.2% of malware samples into their families using approximately 4.6 sec per app. FalDroid can also dramatically reduce the cost of malware investigation by selecting only 8.5% to 22% representative samples that exhibit the most common malicious behavior among all samples. Ming Fan 0002, Jun Liu 0002, Xiapu Luo, Kai Chen 0012, Zhenzhou Tian, Ting Liu 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2018 | Reviving Sequential Program Birthmarking for Multithreaded Software Plagiarism DetectionabstractAs multithreaded programs become increasingly popular, plagiarism of multithreaded programs starts to plague the software industry. Although there has been tremendous progress on software plagiarism detection technology, existing dynamic birthmark approaches are applicable only to sequential programs, due to the fact that thread scheduling nondeterminism severely perturbs birthmark generation and comparison. We propose a framework called TOB (Thread-oblivious dynamic Birthmark) that revives existing techniques so they can be applied to detect plagiarism of multithreaded programs. This is achieved by thread-oblivious algorithms that shield the influence of thread schedules on executions. We have implemented a set of tools collectively called TOB-PD (TOB based Plagiarism Detection tool) by applying TOB to three existing representative dynamic birthmarks, including SCSSB (System Call Short Sequence Birthmark), DYKIS (DYnamic Key Instruction Sequence birthmark) and JB (an API based birthmark for Java). Our experiments conducted on large number of binary programs show that our approach exhibits strong resilience against state-of-the-art semantics-preserving code obfuscation techniques. Comparisons against the three existing tools SCSSB, DYKIS and JB show that the new framework is effective for plagiarism detection of multithreaded programs. The tools, the benchmarks and the experimental results are all publicly available. Zhenzhou Tian, Ting Liu 0002, Eryue Zhuang, Ming Fan 0002, Zijiang Yang 0006 |
IEEE Trans. Software Eng. | 1 |
| 2017 | DAPASA: Detecting Android Piggybacked Apps Through Sensitive Subgraph AnalysisabstractWith the exponential growth of smartphone adoption, malware attacks on smartphones have resulted in serious threats to users, especially those on popular platforms, such as Android. Most Android malware is generated by piggybacking malicious payloads into benign applications (apps), which are called piggybacked apps. In this paper, we propose DAPASA, an approach to detect Android piggybacked apps through sensitive subgraph analysis. Two assumptions are established to reflect the different invocation patterns of sensitive APIs in the injected malicious payloads (rider) of a piggybacked app and in its host app (carrier). With these two assumptions, DAPASA generates a sensitive subgraph (SSG) to profile the most suspicious behavior of an app. Five features are constructed from SSG to depict the invocation patterns. The five features are fed into the machine learning algorithms to detect whether the app is piggybacked or benign. DAPASA is evaluated on a large real-world data set consisting of 2551 piggybacked apps and 44 921 popular benign apps. Extensive evaluation results demonstrate that the proposed approach exhibits an impressive detection performance compared with that of three baseline approaches even with only five numeric features. Furthermore, the proposed approach can complement permission-based approaches and API-based approaches with the combination of our five features from a new perspective of the invocation structure. Ming Fan 0002, Jun Liu 0002, Wei Wang 0012, Haifei Li 0002, Zhenzhou Tian, Ting Liu 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2016 | Frequent Subgraph Based Familial Classification of Android MalwareabstractThe rapid growth of Android malware poses great challenges to anti-malware systems because the sheer number of malware samples overwhelm malware analysis systems. A promising approach for speeding up malware analysis is to classify malware samples into families so that the common features in malwares belonging to the same family can be exploited for malware detection and inspection. However, the accuracy of existing classification solutions is limited because of two reasons. First, since the majority of Android malware is constructed by inserting malicious components into popular apps, the malware's legitimate part may misguide the classification algorithms. Second, the polymorphic variants of Android malware could evade the detection by employing transformation attacks. In this paper, we propose a novel approach that constructs frequent subgraph (fregraph) to represent the common behaviors of malwares in the same family for familial classification of Android malware. Moreover, we propose and develop FalDroid, an automatic system for classifying Android malware according to fregraph, and apply it to 6,565 malware samples from 30 families. The experimental results show that FalDroid can correctly classify 94.5% malwares into their families using around 4.4s per app. Ming Fan 0002, Jun Liu 0002, Xiapu Luo, Kai Chen 0012, Zhenzhou Tian, Xiaodong Zhang 0014, Ting Liu 0002 |
ISSRE | 6 |
| 2016 | Exploiting thread-related system calls for plagiarism detection of multithreaded programs
Zhenzhou Tian, Ting Liu 0002, Ming Fan 0002, Eryue Zhuang, Zijiang Yang 0006 |
J. Syst. Softw. | 1 |
| 2015 | Software Plagiarism Detection with Birthmarks Based on Dynamic Key Instruction SequencesabstractA software birthmark is a unique characteristic of a program. Thus, comparing the birthmarks between the plaintiff and defendant programs provides an effective approach for software plagiarism detection. However, software birthmark generation faces two main challenges: the absence of source code and various code obfuscation techniques that attempt to hide the characteristics of a program. In this paper, we propose a new type of software birthmark called DYnamic Key Instruction Sequence (DYKIS) that can be extracted from an executable without the need for source code. The plagiarism detection algorithm based on our new birthmarks is resilient to both weak obfuscation techniques such as compiler optimizations and strong obfuscation techniques implemented in tools such as SandMark, Allatori and Upx. We have developed a tool called DYKIS-PD (DYKIS Plagiarism Detection tool) and conducted extensive experiments on large number of binary programs. The tool, the benchmarks and the experimental results are all publicly available. Zhenzhou Tian, Ting Liu 0002, Ming Fan 0002, Eryue Zhuang, Zijiang Yang 0006 |
IEEE Trans. Software Eng. | 1 |
| 2014 | Plagiarism detection for multithreaded software based on thread-aware software birthmarksabstractThe availability of inexpensive multicore hardware presents a turning point in software development. In order to benefit from the continued exponential throughput advances in new processors, the software applications must be multithreaded programs. As multithreaded programs become increasingly popular, plagiarism of multithreaded programs starts to plague the software industry. Although there has been tremendous progress on software plagiarism detection technology, existing dynamic approaches remain optimized for sequential programs and cannot be applied to multithreaded programs without significant redesign. This paper fills the gap by presenting two dynamic birthmark based approaches. The first approach extracts key instructions while the second approach extracts system calls. Both approaches consider the effect of thread scheduling on computing software birthmarks. We have implemented a prototype based on the Pin instrumentation framework. Our empirical study shows that the proposed approaches can effectively detect plagiarism of multithread programs and exhibit strong resilience to various semantic-preserving code obfuscations. Zhenzhou Tian, Ting Liu 0002, Ming Fan 0002, Xiaodong Zhang 0014, Zijiang Yang 0006 |
ICPC | 1 |
| 2014 | DBPD: A Dynamic Birthmark-based Software Plagiarism Detection Tool
Zhenzhou Tian, Ming Fan 0002, Eryue Zhuang, Haijun Wang 0002, Ting Liu 0002 |
SEKE | 1 |