VLDB 2026 Research / reviewers in the wild / expert
Dongdong Zhao 0001
dblp:48/2034-1
· DBLP profile ↗
45ranked-venue papers
15as first author
31since 2021 · last 2026
0000-0002-4697-6901ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 4 first-author · 10 since 2021Security and privacy · 12 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningabstractBackdoor attacks pose a persistent security risk to deep neural networks (DNNs) due to their stealth and durability. While recent research has explored leveraging model unlearning mechanisms to enhance backdoor concealment, existing attack strategies still leave persistent traces that may be detected through static analysis. In this work, we introduce the first paradigm of revocable backdoor attacks, where the backdoor can be proactively and thoroughly removed after the attack objective is achieved. We formulate the trigger optimization in revocable backdoor attacks as a bilevel optimization problem: by simulating both backdoor injection and unlearning processes, the trigger generator is optimized to achieve a high attack success rate (ASR) while ensuring that the backdoor can be easily erased through unlearning. To mitigate the optimization conflict between injection and removal objectives, we employ a deterministic partition of poisoning and unlearning samples to reduce sampling-induced variance, and further apply the Projected Conflicting Gradient (PCGrad) technique to resolve the remaining gradient conflicts. Experiments on CIFAR-10 and ImageNet demonstrate that our method maintains ASR comparable to state-of-the-art backdoor attacks, while enabling effective removal of backdoor behavior after unlearning. This work opens a new direction for backdoor attack research and presents new challenges for the security of machine learning systems. Baogang Song, Dongdong Zhao 0001, Jianwen Xiang, Qiben Xu, Zizhuo Yu |
AAAI | 2 |
| 2026 | Industrial Protocol Data Model and State Model Generation Based on Large Language Model
Songsong Liao, Dongdong Zhao 0001, Qianrong Zheng, Junwei Jiang, Jianwen Xiang |
KSEM (1) | 3 |
| 2026 | Generation of Hard SAT Instances and Its Application in Negative Databases for Privacy EnhancementabstractIn recent years, machine learning and deep learning have made remarkable progress and are now widely applied in various fields, including image classification, autonomous driving, natural language processing, and medical diagnosis. However, training these models requires large datasets, which often contain substantial amounts of sensitive personal information, such as medical records and financial details. Without effective privacy protection measures during model training, the risk of sensitive data leakage increases, potentially resulting in severe privacy violations and a loss of trust. As an innovative data representation technique, the Negative Database has proven to be an effective solution in privacy-sensitive domains. Negative databases can be derived from SAT (Boolean Satisfiability Problem) instances, and the hardness of these instances is directly correlated with the level of data protection provided by the Negative Databases. Developing efficient SAT instance generation algorithms to create harder SAT instances can significantly enhance the privacy protection capabilities of the Negative Database. This paper analyzes the hardness conditions of SAT solvers using the Conflict-Driven Clause Learning strategy and proposes a two-stage SAT instance generation algorithm to generate harder SAT instances. These hard instances not only aid in constructing more secure negative databases for enhanced privacy protection but also provide valuable test cases for evaluating and improving SAT solvers. Dongdong Zhao 0001, Pang Chen, Changtian Song, Jianwen Xiang, Junwei Zhou 0002, Zebo Tang, Baogang Song |
IEEE Trans. Big Data | 1 |
| 2026 | BioDeepHash: Generating Consistent Templates for Secure Biometric RecognitionabstractGiven the immutability of biometric data, it is imperative to develop a biometric template protection method that guarantees the complete non-disclosure of any original biometric information while ensuring high recognition performance. Two mainstream approaches in biometric template protection—cancelable biometrics and biometric cryptosystems—have been widely adopted; however, the protected templates produced by these methods still contain some of the original biometric data, which can lead to privacy leakage. To address these challenges, we propose a novel framework named BioDeepHash that integrates deep hashing with cryptographic hash functions. In our approach, a deep hashing model generates consistent templates for similar biometric data from the same user, thereby eliminating intra-class variations. An application-specific XOR string is then used to achieve revocability, and finally these consistent templates are processed by cryptographic hash functions to produce protected templates that meet strict security standards. Our experimental results show that, compared with existing methods, BioDeepHash increases the average Genuine Acceptance Rate by 10.12% on the iris dataset and by 3.12% on the facial dataset, while achieving an extremely low False Acceptance Rate—0% for the iris dataset and only 0.0002% for the facial dataset. Baogang Song, Dongdong Zhao 0001, Jiang Yan, Huanhuan Li 0002, Hao Jiang 0023 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Privacy-Preserving Detection and Defense of Adversarial ExamplesabstractNowadays, deep learning techniques play a crucial role in fields such as computer vision and natural language processing. However, the security and privacy of deep learning models are still threatened. One common type of attack is adversarial attacks, where attackers construct adversarial examples that appear indistinguishable from normal samples to human eyes and induce deep learning models to make misjudgments, potentially leading to catastrophic consequences in deployed systems. Additionally, deep learning models rely heavily on large amounts of training data, once the user information is exposed, the user privacy and security will be compromised. The challenge lies in training deep learning models that are robust against adversarial attacks while protecting user privacy. Most of the existing detection and defense methods are only effective for specific types of attacks and are too dependent on the specific principles of adversarial attacks. Particularly, they do not consider privacy-preserving. In this paper, a privacy-preserving method for adversarial sample detection and defense is proposed, which utilizes denoisers and detectors to defend against adversarial samples while employing differential privacy mechanisms to protect user privacy. Experimental results demonstrate that this method can provide differential privacy protection, adversarial sample detection and defense without significantly affecting the model's accuracy. It effectively defends against common adversarial attack methods such as the Fast Gradient Sign Method (FGSM), Iterative and Projected Gradient Descent (PGD). Furthermore, our method achieves a trade-off between the model's privacy, adversarial robustness, and accuracy. Dongdong Zhao 0001, Guangyue Guo, Xiuwen Lu, Changtian Song |
CSCWD | 1 |
| 2025 | Cancelable iris template based on slicing
Qianrong Zheng, Jianwen Xiang, Changtian Song, Rivalino Matias, Songsong Liao, Dongdong Zhao 0001 |
Comput. Secur. | 9 |
| 2025 | Semi-supervised method for anomaly detection in HTTP trafficabstractAnomaly detection in HTTP traffic is critical for securing web applications against evolving cyber threats. We propose a semi-supervised method that combines domain-specific language modeling with sequence reconstruction to identify anomalies in HTTP requests. Our approach leverages only benign traffic for training and uses reconstruction errors for detecting malicious activity. It achieves a strong balance between precision and recall while maintaining low computational requirements, making it suitable for real-time and edge deployments. Extensive evaluations on three public HTTP datasets show that our method outperforms traditional baselines and fine-tuned BERT models, with an F1-score of 0.92 and AUC of 0.96. We also introduce a simple interpretability mechanism by attributing anomalies to token-level reconstruction errors, providing insights into detected threats. The proposed solution is scalable, lightweight, and effective across diverse attack scenarios without requiring large labeled datasets. Malki Ishara Wasundara, Junwei Zhou 0002, Yanchao Yang 0002, Dongdong Zhao 0001, Jianwen Xiang |
EURASIP J. Inf. Secur. | 4 |
| 2025 | Protected template classification for iris biometrics
Qianrong Zheng, Jianwen Xiang, Songsong Liao, Ling Dong, Dongdong Zhao 0001 |
Expert Syst. Appl. | 6 |
| 2025 | NegSPQ: Similar Patient Query Based on Negative Representation of Genomic DataabstractWith the development of genome sequencing technology, genomic data have been widely collected and used in real-world scenarios, such as genomic medicine and similar patient query (SPQ) cases, based on similarity comparisons of genomic sequences. However, genomic data are unique for every person and contain a large amount of sensitive information involving personal privacy. Therefore, determining how to protect privacy while utilizing genomic data has become a key issue. In this paper, we mainly investigate the privacy protection of SPQs, and we propose five algorithms (NDB-ED, NDB-Band, NDB-Block, NDB-SIS and NDB-SDS) based on a promising technique called negative representation of information (NRI). The proposed algorithms use five kinds of similarity comparison approaches widely applied in SPQs and convert all genomic sequences into negative databases (NDBs, among the most important NRI forms) for privacy protection. When performing an SPQ, NDB-ED approximates the edit distance between the sketches (the statistics of NDBs) of two genomic sequences to evaluate the dissimilarity. Banded edit distance and block-wise edit distance are two effective approximate edit distance algorithms, which can greatly reduce the time complexity. NDB-Band and NDB-Block are used to estimate the banded edit distance and the block-wise edit distance between the sketches of NDB pairs, respectively, to further improve query efficiency and reduce communication overhead. Besides, private genome set intersection size (SIS) and set difference size (SDS) can also be used instead of edit distance to evaluate the dissimilarity between genomic pairs during SPQs. NDB-SIS and NDB-SDS estimate the SIS and SDS between the sketches, for similarity comparison. The experimental results demonstrate that the proposed algorithms can achieve promising results in terms of accuracy and efficiency (Our best algorithm improves accuracy by at least 10% and query time is at least 700 times faster than existing algorithms during SPQs). Dongdong Zhao 0001, Qiben Xu, Yiheng Mao, Jianwen Xiang, Huanhuan Li 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Cross-Project Aging-Related Bug Prediction Based on Transfer Learning and Class Imbalance LearningabstractSoftware aging results from aging-related bugs (ARBs) in long-running systems, that usually causes performance decline and system crashes. Since collecting ARB data is challenging due to its scarcity, it hinders the development of effective prediction models. Moreover, existing cross-project ARB prediction methods often ignore project-specific distribution differences and neglect class imbalance and overlap issues between ARB and non-ARB classes. In this paper, a hybrid approach that combines the balanced distribution adaptation (BDA), the improved subclass discriminant analysis (ISDA), and the self-paced ensemble under-sampling (SPE) techniques, called BISP in short, is proposed to address the aforementioned problems. The main idea behind BISP is first to use BDA to adaptively reduce the difference of projects' marginal distribution and conditional distribution, and then employ ISDA and SPE to alleviate the severe class imbalance together with class overlap. Experimental results obtained for six classifiers and six cross-project datasets show that compared with the state-of-the-art approaches TLAP and JDA-ISDA based on transfer learning, BISP improves the average balance by 34.8% and 2.3% and improves the average AUC by 26.5% and 8.4%, respectively. Compared with the deep learning approach SRLA, BISP can improve the average balance value by 5.1%. Bin Xu 0020, Dongdong Zhao 0001, Junwei Zhou 0002, Wenzhi Xie, Jianwen Xiang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | LogDLR: Unsupervised Cross-System Log Anomaly Detection Through Domain-Invariant Latent RepresentationabstractLog anomaly detection aims to discover abnormal events from massive log data to ensure the security and reliability of software systems. However, due to the heterogeneity of log formats and syntaxes across different systems, existing log anomaly detection methods often need to be designed and trained for specific systems, lacking generalization ability. To address this challenge, we propose LogDLR, a novel unsupervised cross-system log anomaly detection method. The core idea of LogDLR is to use universal sentence embeddings and a Transformer-based autoencoder to extract domain-invariant latent representations from log entries, which can effectively adapt to log format changes and capture semantic information and dependencies in log sequences. To obtain domain-invariant latent representations, we adopt a domain-adversarial training strategy, introducing a domain discriminator that competes with the Transformer-based encoder through a gradient reversal layer, forcing the encoder to learn shared knowledge between different system logs. Finally, the Transformer-based decoder detects anomalies based on the domain-invariant representations obtained by the encoder. We evaluate LogDLR in simulated cross-system scenarios using three publicly available log datasets. The experimental results show that LogDLR can handle heterogeneous logs effectively in cross-system scenarios and achieve efficient and accurate anomaly detection on both source and target systems. Junwei Zhou 0002, Shaowen Ying, Shulan Wang, Dongdong Zhao 0001, Jianwen Xiang, Kaitai Liang, Peng Liu 0005 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Not All Tokens Are Equal: Membership Inference Attacks Against Fine-tuned Language ModelsabstractMembership inference attacks (MIAs), which aim to determine whether a specific sample is included in a machine learning model’s training set, have been recognized as a major privacy threat in recent years. Recent studies have shown that membership inference attacks can lead to privacy leaks in language models. Among existing methods for language model membership inference attacks, reference-based attacks exhibit good performance but require the attacker to obtain the training data distribution of the target model, which is impractical. Reference-free attacks impose lower requirements on attackers but tend to produce unsatisfactory results due to their excessive reliance on the overfitting of the target model. In this paper, we propose a practical Membership Inference Attack based on Weight-enhanced Likelihood (WEL-MIA). We investigate the membership inference difficulty at the token level and find that there are greater mean discrepancy of membership signals between members and non-members for tokens which are more difficult to predict. We also recognize that unreliable calibration probabilities pose an impediment to reference-based membership inference attacks. Anchored on the observation, we design different weights for tokens in the text, providing a new way of aggregating token-level member-ship signals for individual samples. Our attack uses a reference model to calibrate token-level probabilities, with the reference model fine-tuned on the target dataset, thereby producing more reliable calibration probabilities. In our assumption, attackers do not possess any auxiliary data, eliminating the need for the reference model to have prior knowledge of the same domain or distribution. We validate our method on several language models and datasets, and the results demonstrate that WEL-MIA achieves significant performance without relying on any auxiliary data. Changtian Song, Dongdong Zhao 0001, Jianwen Xiang |
ACSAC | 2 |
| 2024 | AEDD: Anomaly Edge Detection Defense for Visual Recognition in Autonomous Vehicle SystemsabstractVisual recognition algorithms based on deep neural network (DNN) have been widely used in the design of automatic driving to recognize traffic sign images. However, there exists adversarial patches which are essentially the anormal image block that can be locally observed but not noticed by humans. And these visual recognition algorithms often suffer from the effect of adversarial patches, due to these patches can change the algorithms recognition result of the images. To solve the above issues, this work proposes the anomaly edge detection and image inpainting defense (AEDD) for visual recognition. This framework uses anomaly location to obtain the anomaly area, uses edge detection to get an accurate edge of the anomaly area, finally uses image inpainting to repair this area. We also combine two attack algorithms with three patch sizes, and generate six types of adversarial patches on the GTSRB dataset. We have demonstrated the effectiveness of our approach, resulting in an average 6.6% increase in defense accuracy compared to the state-of-the-art methods. Our code is available at https://github.com/drtt438/AEDD for the purpose of reproducibility. Junwei Zhou 0002, Dongdong Zhao 0001, Dongqing Liao, Jianwen Xiang |
CSCWD | 3 |
| 2024 | CIDF: Combined Intrusion Detection Framework in Industrial Control Systems based on Packet Signature and Enhanced FSFDPabstractIndustrial Control System (ICS) is vital to critical infrastructures, yet it faces increasing security threats. Current Intrusion Detection System (IDS) designed for ICS often overlooks the unbalanced resource distribution among devices at different layers and primarily focus on known attacks, rendering it difficult to be deployed on all key nodes and vulnerable to unknown threats. To address above issues, we propose a Combined Intrusion Detection Framework (CIDF). This innovative approach is based on strategy of “multi-level layered deployment, combined detection”, deploying the Packet Signature model and the Enhanced Fast Search and Find of Density Peaks (EFSFDP) model on devices at different layers. To achieve optimal use of resource and full protection for ICS and combining the advantages of multiple detection methods to effective detect both known and unknown attacks. The Evaluation using a public gas pipeline dataset and a private dataset shows our approach outperforms existing methods, achieving an average Accuracy, Precision, and Recall of 94%, 95.5%, and 86.5% respectively, and along with superior detection speed. Jianwen Xiang, Qianrong Zheng, Longmin Deng, Dongdong Zhao 0001, Junwei Zhou 0002 |
Internetware | 5 |
| 2024 | A Privacy-Preserving Source Code Vulnerability Detection Method
Dongdong Zhao 0001, Zizhuo Yu, Jianwen Xiang |
PRCV (3) | 1 |
| 2024 | Block-Feature Fusion for Privacy-Protected Iris RecognitionabstractEnsuring privacy often results in sacrificing the accuracy of iris recognition systems. A primary challenge in contemporary iris biometric privacy methods lies in striking a balance between recognition accuracy and privacy of protected iris templates. Hence, any proposition for privacy-protected iris recognition must prioritize irreversibility, revocability, and unlinkability to uphold robust privacy standards while achieving higher recognition accuracy. This research proposes an approach that stands as a robust solution with acceptable advancement in three challenges: recognition accuracy, privacy protection and computational efficiency. We experimented with an innovative technique that manipulates the columns of an iris template by fusing the bit pattern of the template using the XOR operation. The transformation process is non-linear. This fusion introduces randomness and variability to the fused templates. It also poses enhanced privacy protection. Two datasets were used to validate the proposed approach. Based on the results of dataset 1, the proposed approach accepts genuine users at a rate of 99.11% while it accepts 0.01% of imposters. For dataset 2, the Genuine Acceptance Rate (GAR) is depicted as 81.12% while FAR is at 0.01%. The proposed approach can be applied in practice due to its higher computational efficiency. As further improvements, the research can be extended to more widespread databases and higher-quality iris samples. Wiraj Udara Wickramaarachchi, Junwei Zhou 0002, Dongdong Zhao 0001, Jianwen Xiang |
TrustCom | 3 |
| 2024 | An effective iris biometric privacy protection scheme with renewability
Wiraj Udara Wickramaarachchi, Dongdong Zhao 0001, Junwei Zhou 0002, Jianwen Xiang |
J. Inf. Secur. Appl. | 2 |
| 2024 | A cancellable iris template protection scheme based on inverse merger and Bloom filter
Qianrong Zheng, Jianwen Xiang, Songsong Liao, Dongdong Zhao 0001 |
J. Inf. Secur. Appl. | 6 |
| 2024 | PMTT: Parallel multi-scale temporal convolution network and transformer for predicting the time to aging failure of software systems
Xiao Yu 0008, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang |
J. Syst. Softw. | 5 |
| 2024 | TTAFPred: Prediction of time to aging failure for software systems based on a two-stream multi-scale features fusion network
Xiao Yu 0008, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang |
Softw. Qual. J. | 5 |
| 2023 | DLMT: Outsourcing Deep Learning with Privacy Protection Based on Matrix TransformationabstractIn recent years, deep learning has been applied in a wide variety of domains and gains outstanding success. In order to achieve high accuracy, a large amount of training data and high-performance hardware are necessary for deep learning. In real-world applications, many deep learning developers usually rent cloud GPU servers to train or deploy their models. Since training data may contain sensitive information, training models on cloud servers will cause severe privacy leakage problem. To solve this problem, we propose a privacy-preserving deep learning model based on matrix transformation. Specifically, we transform original data by adding or multiplying a random matrix. The obtained data is significantly different from the origin and it is hard to recover original data, so it can protect the privacy in original data. Experimental results demonstrate that the models trained with processed data can achieve high accuracy. Dongdong Zhao 0001, Jianwen Xiang, Huanhuan Li 0002 |
CSCWD | 1 |
| 2023 | IFCM: An improved Fuzzy C-means clustering method to handle Class Overlap on Aging-related Software Bug PredictionabstractSoftware aging refers to a problem of performance decay in long-running software systems. This phenomenon is primarily attributed to the accumulation of run-time errors, commonly known as aging-related bugs (ARBs). Detecting ARBs through Aging-related Bug Prediction (ARBP) is crucial in ensuring system reliability. The effectiveness of ARBP heavily relies on the quality of datasets. However, ARB datasets often suffer from class overlap, where instances from different classes exhibit similar feature values. Class overlap poses a significant challenge as it compromises the quality of training data and subsequently impacts ARBP accuracy. To address this issue, we propose an improved Fuzzy C-means clustering method named IFCM, designed to mitigate class overlap in ARBP tasks. IFCM can identify whether an instance occurs overlap, and identify the overlap degree of this instance through the predefined parameters. We evaluate our proposed method on two public datasets Linux and MySQL and one self-collected dataset NetBSD using five different classifiers with five performance metrics (AUC, F1, Balance, PD, PF). Comparison with four existing methods (No clean, NCL, IKMCCA, ROCT) demonstrates that IFCM is effective in alleviating class overlap in ARBP. For Instance, IFCM achieves promising results in terms of AUC blue (which are 0.762, 0.757, and 0.642) and Balance (which are 0.709, 0.736, and 0.595) at the dataset level. Shuo Feng 0003, Wenzhi Xie, Dongdong Zhao 0001, Jianwen Xiang, Roberto Pietrantuono, Roberto Natella, Domenico Cotroneo |
ISSRE | 4 |
| 2023 | Cancelable Iris Biometrics Based on Transformation NetworkabstractThe application of iris biometric data has become prevalent across various domains, encompassing access control, identity verification, and criminal investigations. Consequently, there is a pressing need to develop effective methods for safeguarding the privacy of iris data. While numerous methods for iris data protection have been proposed, the majority of them fall short of meeting the ISO/IEC 24745 standards about irreversibility, revocability, and unlinkability. In this paper, we introduce a novel iris data protection method called TNCB, which is based on a transformation network. The TNCB involves performing a block-wise permutation of the original iris images using application-specific parameters, followed by pixel-by-pixel modulo and inversion fusion operations. The resulting images are subsequently employed for pre-training a recognition network that will be used to recognize protected images. Afterwards, a transformation network is introduced to achieve a further non-invertible transformation. Our security analysis demonstrates that the TNCB could fulfill the three major security requirements. To validate its effectiveness, we conducted a series of attack and performance experiments on the CASIA-Iris-Lamp and CASIA-Iris-Thousand datasets. Experimental results substantiated the robustness of TNCB in maintaining recognition performance while safeguarding the privacy of iris data. Furthermore, experimental results also highlight that our scheme could effectively support iris recognition in both open-set and close-set modes. Dongdong Zhao 0001, Hucheng Liao, Songsong Liao, Huanhuan Li 0002, Jianwen Xiang |
QRS | 1 |
| 2023 | Generating Random SAT Instances: Multiple Solutions could be Predefined and Deeply HiddenabstractThe generation of SAT instances is an important issue in computer science, and it is useful for researchers to verify the effectiveness of SAT solvers. Addressing this issue could inspire researchers to propose new search strategies. SAT problems exist in various real-world applications, some of which have more than one solution. However, although several algorithms for generating random SAT instances have been proposed, few can be used to generate hard instances that have multiple predefined solutions. In this paper, we propose the KHidden-M algorithm to generate SAT instances with multiple predefined solutions that could be hard to solve by the local search strategy when the number of predefined solutions is small enough and the Hamming distance between them is not less than half of the solution length. Specifically, first, we generate an SAT instance that is satisfied by all of the predefined solutions. Next, if the generated SAT instance does not satisfy the hardness condition, then a strategy will be conducted to adjust clauses through multiple iterations to improve the hardness of the whole instance. We propose three strategies to generate the SAT instance in the first part. The first strategy is called the random strategy, which randomly generates clauses that are satisfied by all of the predefined solutions. The other two strategies are called the estimating strategy and greedy strategy, and using them, we attempt to generate an instance that directly satisfies or is closer to the hardness condition for the local search strategy. We employ two SAT solvers (i.e., WalkSAT and Kissat) to investigate the hardness of the SAT instances generated by our algorithm in the experiments. The experimental results show the effectiveness of the random, estimating and greedy strategies. Compared to the state-of-the-art algorithm for generating SAT instances with predefined solutions, namely, M-hidden, our algorithm could be more effective in generating hard SAT instances. Dongdong Zhao 0001, Wenjian Luo, Jianwen Xiang, Hao Jiang 0023 |
J. Artif. Intell. Res. | 1 |
| 2022 | Negative Survey with Aggregate Scores from Multiple QuestionsabstractNegative survey is a sensitive data collection method with a wide range of application scenarios. Contrary to traditional surveys, participants are asked to randomly choose an option which they do not belong to, and thus, their privacy can be protected. After collecting the data from participants, the overall distribution of participants over different options can be obtained through reconstruction algorithms. However, existing negative survey models have not considered the questionnaire which has multiple questions and should compute aggregate scores to make overall evaluation. Thus, this paper proposes to retain aggregate scores in negative surveys, and proposes an algorithm to exploit the aggregate scores during reconstructing results to enhance the accuracy. Experimental results demonstrate that the proposed approach could outperform existing algorithms. Dongdong Zhao 0001 |
CSCWD | 2 |
| 2022 | The Impact of Software Aging and Rejuvenation on the User Experience for Android SystemabstractIn the Android system, software aging is an essential factor affecting user experience. Its occurrence will lead to poor responsiveness or crash/hang failure of the system. Recently, the strategies to schedule rejuvenation are marching toward a situation that needs to consider both usage behavioral aspects of its users (i.e., switch between active and sleep modes) and two-level software aging process (i.e., Operating System (OS) and Application Software (AS)), because rejuvenating the OS or AS during active time slot contributes to terrible user experience. To be able to achieve higher user experience and lower user interference, in this paper, we present to employ the Continuous Time Markov Chain (CTMC) model to study the impact of software aging and rejuvenation on user experience on two different rejuvenation strategies: condition-based and time-based rejuvenations. In contrast to the existing works, our models capture the interactions between usage behavioral aspects of users and two-level aging and rejuvenation. We then define three metrics to evaluate the user experience, including User-perceived (1) Fluency (UF), (2) Failure Probability (UFP), and (3) Availability (UA). The numerical analysis has the following noticed conclusions. The optimal value of UF yielded by condition-based rejuvenation reaches a 3.486% improvement over that of time-based. Therefore, the former is an appealing rejuvenation solution. Moreover, compared with single-level (OS and AS) rejuvenation models, two-level rejuvenation indeed improves the user experience. Concretely, the values of three metrics achieve 80.20% and 14.45%,83.39% and 98.45%, 0.004% and 0.048% improvements, respectively. Xiao Yu 0008, Dongdong Zhao 0001, Jianwen Xiang |
ISSRE | 5 |
| 2022 | CBSDI: Cross-Architecture Binary Code Similarity Detection based on Index TableabstractBinary code similarity detection for cross-platform is widely used in plagiarism detection, malware detection and vulnerability search, aiming to detect whether two binary functions over different platforms are similar. Existing cross-architecture approaches mainly rely on the approximate matching calculation of complex high-dimensional features, such as graph, which are inevitably slow and unsuitable for large-scale applications. To solve this problem, we propose a novel approach based on index table called CBSDI, improving efficiency by screening a batch of mismatched functions before similarity detection. We select three features and compare them across architectures to select the most appropriate one to construct the index table, and this table can be embedded in other tools. The evaluation shows that the index table can roughly cut the computational costs in half when there are few errors. Moreover, compared with the related works in the literature, our proposed approach can improve not only the efficiency but also the accuracy. Longmin Deng, Dongdong Zhao 0001, Junwei Zhou 0002, Zhe Xia, Jianwen Xiang |
QRS | 2 |
| 2022 | GAN-Based Privacy-Preserving Unsupervised Domain AdaptationabstractIn recent years, the rapid development of deep learning is attributed to the large amount of labeled data brought by the digital age. When there is no labeled data available in some application scenarios, domain adaptation can be used to transfer knowledge from the source domain with labeled data to the target domain without labeled data. In the process of domain adaptation, the target client requires direct access to the source data or model, which would lead to the risk of privacy leakage, e.g., Membership Inference Attacks (MIA). Attackers can collect the prediction vector of the model through black-box access to the source model, and then infer an individual’s membership in the source training dataset. To deal with this privacy issue, we propose a GAN-based Privacy-Preserving Unsupervised Domain Adaptation Framework. Specifically, the target client learns a conditional generator, sends the intermediate results perturbed by differential privacy to the source client, and the source client uses the source model to provide guidance for the generator so that the generator can generate the data corresponding to the input label that is the same as the data distribution in the target domain. We evaluate the performance of our proposed method on digital dataset and office-31dataset, which are popular domain adaptation benchmark datasets, and verify the security by the accuracy and F1-score of Membership Inference Attacks. Dongdong Zhao 0001, Huanhuan Li 0002, Jianwen Xiang |
QRS | 1 |
| 2021 | An Efficient Approximation for Quantitative Analysis of Dynamic Fault TreesabstractThis paper presents a feasibility and effective ap-proximation method to estimate the failure probability of the top event of a dynamic fault tree. The method is based on a minimal canonical form and uses a quantitative relationship between the smallest cut sequence and the entire sequence. Comparison with discrete-time Bayesian networks and Monte Carlo simulation methods, the validity of this method is assessed on two case studies approximating the probabilities of the top event of a Hypothetical Cardiac Assist System (HCAS) and a fictitious system. The case study results show that our method can achieve similar accuracy with smaller relative error and shorter execution time. Luyao Ye, Erqing Li, Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
ISSRE | 3 |
| 2021 | Reliability-redundancy allocation problem considering imperfect fault coverageabstractThe reliability-redundancy allocation problem (RRAP) has been widely investigated during the last decade. In most of existing studies, component failures are assumed to be covered perfectly which means all faults can be timely detected, located, and isolated. However, the coverage could be imperfect in reality and a not-covered component failure may lead to system failure without constraint. In this paper, the RRAP is solved considering the imperfect fault coverage model (IFCM, only faulty components can be covered) and the irrelevance coverage model (ICM, both faulty and irrelevant components can be covered). It has been proved that an excessive level of redundancy may reduce the system reliability rather than improve it when the fault coverage is imperfect. Therefore, when the IFCM and the ICM are considered in the RRAP, in addition to resource constraints, the coverage model itself also limits the level of redundancy. Three benchmark problems are investigated in this paper. The genetic algorithm is adopted to solve the new mixed integer nonlinear programming problem. The results show that the optimal designs of system in the two coverage models are different from the existing researches that only consider the perfect fault coverage model. The redundant components used in the optimal solution are less than the existing studies. The advantage of the ICM over the IFCM is also verified in this paper. Dongdong Zhao 0001, Jianwen Xiang |
QRS | 3 |
| 2021 | Quantitative Analysis of the Dynamic Relevance of SystemsabstractIn systems with imperfect fault coverage (IFC), all components are subject to uncovered failures, possibly threatening the whole system. Therefore, to improve the system reliability, it is important to timely detect, identify, and shut down the components that are no more relevant for the system operation. This article addresses quantitative evaluation of the relevance of components, assuming that they have independent and identically distributed lifetimes to characterize the impact of the system design only on the system reliability and energy consumption. To this end, the dynamic relevance measure is defined to characterize the irrelevant components in different stages of the system lifetime depending on the number of occurred component failures, supporting the evaluation of the probability that the system fails due to uncovered failures of irrelevant components. Moreover, the system reliability over time is also efficiently derived, both in the case that irrelevance is not considered and in the case that irrelevant components can be immediately isolated, notably supporting any general (i.e., non-Markovian) distribution for the failure time of components. Feasibility and effectiveness of the approach are assessed on two real-scale case studies addressing reliability evaluation of a flight control system and a multihop wireless sensor network. Luyao Ye, Dongdong Zhao 0001, Jianwen Xiang, Laura Carnevali, Enrico Vicario |
IEEE Trans. Reliab. | 2 |
| 2020 | Cross-Project Aging-Related Bug Prediction Based on Joint Distribution Adaptation and Improved Subclass Discriminant AnalysisabstractSoftware aging, which is caused by Aging-Related Bugs (ARBs), refers to the phenomenon of performance degradation and eventual crash in long running systems. In order to discover and remove ARBs, ARB prediction is proposed. However, due to the low presence and reproducing difficulty of ARBs, it is usually difficult to collect sufficient ARB data within a project. Therefore, cross-project ARB prediction is proposed as a solution to build the target project's ARB predictor by using the labeled data from the source project. A key point for cross-project ARB prediction is to reduce distribution difference between source and target project. However, existing approaches mainly focus on the marginal distribution difference while somehow overlook the conditional distribution difference, and they mainly use random oversampling to alleviate the class imbalance which may lead to overfitting. To address these problems, we propose a new crossproject ARB prediction approach based on Joint Distribution Adaptation (JDA) and Improved Subclass Discriminant Analysis (ISDA), called JDA-ISDA. The key idea of JDA-ISDA is first to use JDA to reduce the marginal distribution and conditional distribution difference jointly and then apply ISDA to alleviate the severe class imbalance problem. A set of experiments are carried out on two large open-source projects with six different machine learning (ML) classifiers. The experimental results demonstrate that compared with the state-of-the-art Transfer Learning based Aging-related bug Prediction (TLAP) and Supervised Representation Learning Approach (SRLA), JDA-ISDA is much more robust to different ML classifiers than TLAP, and the average improvement in terms of the balance value can be achieved up to 31.8%, and JDA-ISDA also outperforms TLAP and SRLA on average when logistic regression is chosen as the classifier for best performance prediction. Bin Xu 0020, Dongdong Zhao 0001, Junwei Zhou 0002, Jianwen Xiang |
ISSRE | 2 |
| 2020 | Software aging and rejuvenation in android: new models and metrics
Jianwen Xiang, Caisheng Weng, Dongdong Zhao 0001, Artur Andrzejak 0001, Shengwu Xiong 0001, Lin Li 0001 |
Softw. Qual. J. | 3 |
| 2019 | Reliability Analysis of Phased-Mission System in Irrelevancy Coverage ModelabstractIn a phased-mission system (PMS), an uncovered component fault may lead to a mission failure regardless of the status of other components, and the reliability can be analyzed with traditional imperfect fault coverage model (IFCM). The IFCM, however, only considers the coverage of faulty components. Recently, an irrelevancy coverage model (ICM) is proposed to cover both faulty components and irrelevant components, but the analysis is limited to normal non-phased mission systems. This paper first demonstrates that, the coverage of irrelevant components is also important in PMSs, as an initially relevant component could also become irrelevant later due to the failures of other components, and an uncovered fault of irrelevant component may threaten the whole mission as well. A method to analyze the reliability of PMS in ICM is proposed using sum of disjoint products (SDP) technique. Experimental results demonstrate not only the effectiveness of the proposed reliability analysis method, but also that the ICM can achieve higher reliability than the IFCM for PMSs in general. Dongdong Zhao 0001, Luyao Ye, Jianwen Xiang |
QRS | 2 |
| 2019 | CVSkSA: cross-architecture vulnerability search in firmware based on kNN-SVM and attributed control flow graphabstractTo prevent the same known vulnerabilities from affecting different firmware, searching known vulnerabilities in binary firmware across different architectures is crucial. Because the accuracy of existing cross-architecture vulnerability search methods is not high, we propose a staged approach based on support vector machine (SVM) and attributed control flow graph (ACFG) at the function level to improve the accuracy using prior knowledge. Furthermore, for efficiency, we utilize the k-nearest neighbor (kNN) algorithm to prune and SVM to refine in the function prefilter stage. Although the accuracy of the proposed method using kNN-SVM approach is slightly lower than the accuracy of the method using only SVM, its efficiency is significantly enhanced. We have implemented our approach CVSkSA to search several vulnerabilities in real-world firmware images. The experimental results show that the accuracy of the proposed method using kNN-SVM approach is close to the accuracy of the method using only SVM in most cases, while the former is approximately four times faster than the latter. Dongdong Zhao 0001, Hong Lin 0004, Linjun Ran, Mushuai Han, Shengwu Xiong 0001, Jianwen Xiang |
Softw. Qual. J. | 1 |
| 2018 | Handling Unreasonable Data in Negative Surveys
Jianwen Xiang, Shu Fang, Dongdong Zhao 0001, Shengwu Xiong 0001, Chunhui Yang |
DASFAA (2) | 3 |
| 2018 | Privacy-Preserving K-Means Clustering Upon Negative Databases
Dongdong Zhao 0001, Jianwen Xiang, Xing Liu 0002, Haiying Zhou, Shengwu Xiong 0001 |
ICONIP (4) | 3 |
| 2018 | Iris Template Protection Based on Randomized Response Technique and Aggregated Block InformationabstractNowadays, biometric recognition has been widely used in real-world applications, but it has also brought potential privacy threats to users. Iris template protection enables an effective iris recognition while protecting personal privacy. In this paper, we propose a method for iris template protection based on randomized response technique and aggregated block information. Specifically, the iris data are first permuted according to an application-specific parameter; next, the permuted data are flipped using the randomized response technique; finally, the result is divided into blocks, and the aggregated information (i.e., the sum of all bits) in each block is calculated and stored instead of original iris data for privacy protection. We demonstrate that the proposed method supports the shifting and masking strategies for enhancing recognition performance. Moreover, the proposed method satisfies the three privacy requirements prescribed in ISO/IEC 24745: irreversibility, revocability and unlinkability. Experimental results show that the proposed method could effectively maintain the recognition performance (w.r.t. the original iris recognition system without privacy protection) on the iris database CASIA-IrisV3-Interval. Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
ISSRE | 1 |
| 2018 | Iris Template Protection Based on Local RankingabstractBiometrics have been widely studied in recent years, and they are increasingly employed in real-world applications. Meanwhile, a number of potential threats to the privacy of biometric data arise. Iris template protection demands that the privacy of iris data should be protected when performing iris recognition. According to the international standard ISO/IEC 24745, iris template protection should satisfy the irreversibility, revocability, and unlinkability. However, existing works about iris template protection demonstrate that it is difficult to satisfy the three privacy requirements simultaneously while supporting effective iris recognition. In this paper, we propose an iris template protection method based on local ranking. Specifically, the iris data are first XORed (Exclusive OR operation) with an application-specific string; next, we divide the results into blocks and then partition the blocks into groups. The blocks in each group are ranked according to their decimal values, and original blocks are transformed to their rank values for storage. We also extend the basic method to support the shifting strategy and masking strategy, which are two important strategies for iris recognition. We demonstrate that the proposed method satisfies the irreversibility, revocability, and unlinkability. Experimental results on typical iris datasets (i.e., CASIA-IrisV3-Interval, CASIA-IrisV4-Lamp, UBIRIS-V1-S1, and MMU-V1) show that the proposed method could maintain the recognition performance while protecting the privacy of iris data. Dongdong Zhao 0001, Shu Fang, Jianwen Xiang, Shengwu Xiong 0001 |
Secur. Commun. Networks | 1 |
| 2018 | Negative Iris RecognitionabstractElements of a person's biometrics are typically stable over the duration of a lifetime, and thus, it is highly important to protect biometric data while supporting recognition (it is also called secure biometric recognition). However, the biometric data that are derived from a person usually vary slightly due to a variety of reasons, such as distortion during picture capture, and it is difficult to use traditional techniques, such as classical encryption algorithms, in secure biometric recognition. The negative database (NDB) is a new technique for privacy preservation. Reversing the NDB has been demonstrated to be an NP-hard problem, and several algorithms for generating hard-to-reverse NDBs have been proposed. In this paper, first, we propose negative iris recognition, which is a novel secure iris recognition scheme that is based on the NDB. We show that negative iris recognition supports several important strategies in iris recognition, e.g., shifting and masking. Next, we analyze the security and efficiency of negative iris recognition. Experimental results show that negative iris recognition is an effective and secure iris recognition scheme. Specifically, negative iris recognition can achieve a highly promising recognition performance (i.e., GAR = 98.94% at FAR = 0.01%, EER = 0.60%) on the typical database CASIA-IrisV3-Interval. Dongdong Zhao 0001, Wenjian Luo, Lihua Yue |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | A Novel Negative Location Collection Method for Finding Aggregated LocationsabstractCurrently, many intelligent transportation systems (ITSs) require aggregation of location information. Although users enjoy the convenience of ITSs, privacy concerns have resulted in user's caution in offering location information. Certain studies have been performed to preserve user location privacy. However, most studies that focus on preserving location privacy require a trusted third party or do not consider the movements of users. In this paper, we propose a method based on negative surveys that can be used to estimate the number of people in geographic locations. This method, which can preserve user privacy regardless of user movements, adopts a simple negative survey algorithm for user devices and an estimation algorithm for the server and might be suitable for low-power mobile devices. The experimental results demonstrate that our method is capable of locating the people's gathering places with fine control granularity, which renders it a promising application. Hao Jiang 0023, Wenjian Luo, Dongdong Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | An Ontology-based Knowledge Management System for Software TestingabstractSoftware testing is an important activity in quality assurance and it generates large amount of knowledge.Software testers need to gather domain knowledge to be able to successfully conduct a software testing activity.Not having a proper knowledge base within its own context by software testing environments cause software testers to query limited knowledge available or consult peer software testers, which would greatly impact on their decision-making process.Ontologies emerge as one of the more appropriate knowledge management tools for supporting knowledge representation, processing, storage and retrieval.Given great importance to knowledge for software testing, and the potential benefits of managing software testing knowledge, using semantic web technologies, ontology based knowledge management system is developed.A Software testing knowledge sharing ontology is designed to describe software testing domain knowledge.SPARQL is used as the query language to retrieve software testing knowledge from the semantic storage.Both Ontology experts and non-experts evaluated the developed ontology.We believe our software testing ontology can support other software organizations to improve the sharing of knowledge and learning practices. Shanmuganathan Vasanthapriyan, Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
SEKE | 3 |
| 2017 | One-time password authentication scheme based on the negative database
Dongdong Zhao 0001, Wenjian Luo |
Eng. Appl. Artif. Intell. | 1 |
| 2017 | Experimental analyses of the K-hidden algorithm
Dongdong Zhao 0001, Wenjian Luo, Lihua Yue |
Eng. Appl. Artif. Intell. | 1 |
| 2013 | A Study of the Private Set Intersection Protocol Based on Negative DatabasesabstractNowadays, data privacy has been widely concerned. The private set intersection means that several parties calculate the intersection of their private sets while without revealing extra information about their private data. The negative database (NDB) is a new technique for preserving privacy, and it stores information in the complementary set of a traditional database (DB). Reversing the NDB to recover the corresponding DB is an NP-hard problem, and this property is the security foundation of the NDB. Moreover, the NDB can directly support some database operations such as intersection, union, select and Cartesian product. However, so far, there is no research work about the secure multi-party computation based on NDBs. In this paper, firstly, a two-party private set intersection protocol based on NDBs is proposed, and its security and efficiency are analyzed. Then, the multi-party private set intersection protocol based on NDBs is given. Dongdong Zhao 0001, Wenjian Luo |
DASC | 1 |