VLDB 2026 Research / reviewers in the wild / expert
Ping Luo 0004
dblp:54/4989-4
· DBLP profile ↗
17ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0001-6171-3811ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 3 since 2021Security and privacy · 5 · 4 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DSFM: Enhancing Functional Code Clone Detection with Deep Subtree InteractionsabstractFunctional code clone detection is important for software maintenance. In recent years, deep learning techniques are introduced to improve the performance of functional code clone detectors. By representing each code snippet as a vector containing its program semantics, syntactically dissimilar functional clones are detected. However, existing deep learning-based approaches attach too much importance to code feature learning, hoping to project all recognizable knowledge of a code snippet into a single vector. We argue that these deep learning-based approaches can be enhanced by considering the characteristics of syntactic code clone detection, where we need to compare the contents of the source code (e.g., intersection of tokens, similar flow graphs, and similar subtrees) to obtain code clones. In this paper, we propose a novel deep learning-based approach named DSFM, which incorporates comparisons between code snippets for detecting functional code clones. Specifically, we improve the typical deep clone detectors with deep subtree interactions that compare every two subtrees extracted abstract syntax trees (ASTs) of two code snippets, thereby introducing more fine-grained semantic similarity. By conducting extensive experiments on three widely-used datasets, GCJ, OJClone, and BigCloneBench, we demonstrate the great potential of deep subtree interactions in code clone detection task. The proposed DSFM outperforms the state-of-the-art approaches, including two traditional approaches, two unsupervised and four supervised deep learning-based baselines. Shaohua Qiang, Dinghong Song, Min Zhou 0001, Hai Wan, Xibin Zhao, Ping Luo 0004, Hongyu Zhang 0002 |
ICSE | 7 |
| 2022 | LibDB: An Effective and Efficient Framework for Detecting Third-Party Libraries in BinariesabstractThird-party libraries (TPLs) are reused frequently in software applications for reducing development cost. However, they could introduce security risks as well. Many TPL detection methods have been proposed to detect TPL reuse in Android bytecode or in source code. This paper focuses on detecting TPL reuse in binary code, which is a more challenging task. For a detection target in binary form, libraries may be compiled and linked to separate dynamic-link files or built into a fused binary that contains multiple libraries and project-specific code. This could result in fewer available code features and lower the effectiveness of feature engineering. In this paper, we propose a binary TPL reuse detection framework, LibDB, which can effectively and efficiently detect imported TPLs even in stripped and fused binaries. In addition to the basic and coarse-grained features (string literals and exported function names), LibDB utilizes function contents as a new type of feature. It embeds all functions in a binary file to low-dimensional representations with a trained neural network. It further adopts a function call graph-based comparison method to improve the accuracy of the detection. LibDB is able to support version identification of TPLs contained in the detection target, which is not considered by existing detection methods. To evaluate the performance of LibDB, we construct three datasets for binary-based TPL reuse detection. Our experimental results show that LibDB is more accurate and efficient than state-of-the-art tools on the binary TPL detection task and the version identification task. Our datasets and source code used in this work are anonymously available at https://github.com/DeepSoftwareAnalytics/LibDB. Yanlin Wang 0001, Hongyu Zhang 0002, Shi Han, Ping Luo 0004, Dongmei Zhang 0001 |
MSR | 5 |
| 2022 | Obfuscated code is identifiable by a token-based code clone detection technique
Junaid Akram, Danish Vasan, Ping Luo 0004 |
Int. J. Inf. Comput. Secur. | 3 |
| 2021 | DroidMD: an efficient and scalable Android malware detection approach at source code levelabstractSecurity researchers and anti-virus industries have speckled stress on an Android malware, which can actually damage your phones and threatens the Android markets. In this paper, we propose and develop DroidMD, a scalable self-improvement based tool, based on auto optimisation of signature set, which detect malicious apps in the market at source code level. A prototype has been developed tested and implemented to detect malware in applications. We implement and evaluate our approach on almost 30,000 applications including 27,000 benign and 3,670 malware applications. DroidMD detects malware in different applications at partial level and full level. It analyses only the applications code, which increase its reliability. Our evaluation of DroidMD demonstrates that our approach is very efficient in detecting malware at large scale with high accuracy of 95.5%. Junaid Akram, Majid Mumtaz, Gul Jabeen, Ping Luo 0004 |
Int. J. Inf. Comput. Secur. | 4 |
| 2021 | Vulnerability severity prediction model for software based on Markov chainabstractSoftware vulnerabilities primarily constitute security risks. Commonalities between faults and vulnerabilities prompt developers to utilise traditional fault prediction models and metrics for vulnerability prediction. Although traditional models can predict the number of vulnerabilities and their occurrence time, they fail to accurately determine the seriousness of vulnerabilities, impacts, and severity level. To address these deficits, we propose a method for predicting software vulnerabilities based on a Markov chain model, which offers a more comprehensive descriptive model with the potential to accurately predict vulnerability type, i.e., the seriousness of the vulnerabilities. The experiments are performed using real vulnerability data of three types of popular software: Windows 10, Adobe Flash Player and Firefox. Our model is shown to produce accurate predictive results. Gul Jabeen, Ping Luo 0004 |
Int. J. Inf. Comput. Secur. | 3 |
| 2021 | An improved cryptanalysis of large RSA decryption exponent with constrained secret keyabstractIn this study, we revisit the RSA public key cryptosystem in some special case of Boneh and Durfee's attack when the private key d assumes to be larger than the public key e. The attack in this study is the variation of an approach adopted by Luo et al. (2009) based on large decryption exponent. They had chosen a large private key (d > e) and found the weak keys in some specific range between N0.258 ≤ e ≤ N0.857. We highlight the shortcomings and new improvements in our study with more refined bound analysis up to the range between N0.104 ≤ e ≤ N0.923. Our experimental results revealed more refined bounds using lattice-based Coppersmith's method. In our experimental yield, we find the small roots of the devised polynomial, which helps to factorise the RSA modulus of size up to 1,024-bits. We also measure the probability of a specific range of weak keys, which further certify our results about weak keys in an RSA constrained secret key environment. Majid Mumtaz, Ping Luo 0004 |
Int. J. Inf. Comput. Secur. | 2 |
| 2021 | SQVDT: A scalable quantitative vulnerability detection technique for source code security assessmentabstractSummary Vulnerability detection and exploit is becoming a very important part of security, especially in malware code delivery, hacking a system, efforts to create patches, improving the source code, or updating a software. Vulnerabilities in applications, including browsers, media players, online services, document readers, and so forth. are often exploited and cause a serious damage. In this article, we propose a vulnerability detection technique to detect vulnerabilities in software, as well as shared libraries at source code level. We crawl the vulnerable source code by tracing and locating the patch files from different web sources according to their CVE‐numbers and built a fingerprint index of 2931 vulnerable files. Then we developed a vulnerability detection approach based on code clone detection technique and detect hundreds of vulnerabilities in thousands of GitHub open source projects, which are not noticed before as vulnerable. We detected vulnerabilities in some very famous recently available software, including latest version of Linux, HTC‐kernel, FindX‐8.1‐kernel, and in 7‐TB of C/C++ source code (152,823 open source projects). In this study, we discuss some of the very high severity level (CVSS) vulnerabilities that are detected by our approach. Furthermore, we performed an empirical evaluation and verification on these vulnerabilities, including intraproject clone vulnerabilities, copied‐kernel clone vulnerabilities, and library‐used clone vulnerabilities. Our technique is very fast, efficient, reliable, practical, scalable, and can be implemented at industrial level. The comparison with the state‐of‐the‐art tools shows the effectiveness of our approach. Junaid Akram, Ping Luo 0004 |
Softw. Pract. Exp. | 2 |
| 2020 | OSLDetector: Identifying Open-Source Libraries through Binary AnalysisabstractUsing open-source libraries can provide rich functions and reduce development cost. However, some critical issues have also been caused such as license conflicts and vulnerability risks. In this paper, we design and implement an open-source libraries detection tool OSLDetector which uses methods of matching features to detect third-party libraries for multi-platform software in binaries. We took a series of methods such as filtering features and novelty building an internal clone forest to cope with the challenge of feature duplication. The tool can also provide the conflict of licenses and identify possible corresponding vulnerabilities, so these potential risks can be resolved and avoided. To evaluate the efficiency of OSLDetector, we collect 5K libraries containing 9K versions and manage their respective license type and existing vulnerabilities. The experimental results with a precision of 96% and recall of 92.3% show that OSLDetector is effective and outperforms similar tools. Ping Luo 0004, Min Zhou 0001 |
ASE | 2 |
| 2020 | How to build a vulnerability benchmark to overcome cyber security attacksabstractCybercrimes are on a dramatic rise worldwide. The crime rate is growing day by day in every field or department which is directly or indirectly connected to the internet including Government, business or any individual. The main objective of this study is to evaluate the vulnerabilities in different software systems at the source code level by tracing their patch files. The authors have collected the source code of different types of vulnerabilities at a different level of granularities. They have proposed different ways to collect or trace the vulnerability code, which can be very helpful for security experts, organisations and software developers to maintain security measures. By following their proposed method, you can build your own vulnerability data‐set and can detect vulnerabilities in any system by using suitable code clone detection technique. The study also includes a discussion of reasons for the rise in cybercrimes including zero‐day exploits. A case study has been discussed with results and research questions to show the effectiveness of this study. This study concludes with the effective key findings of published and non‐published vulnerabilities and the ways to prevent from different security attacks to overcome cybercrimes. Junaid Akram, Ping Luo 0004 |
IET Inf. Secur. | 2 |
| 2020 | Remarks on the cryptanalysis of common prime RSA for IoT constrained low power devices
Majid Mumtaz, Ping Luo 0004 |
Inf. Sci. | 2 |
| 2020 | IBFET: Index-based features extraction technique for scalable code clone detection at file level granularityabstractSummary Many techniques have been developed over the years to detect code clones in different software systems to maintain security measures. These techniques often require the source code to compare the subject system against a very large data set of big code. This paper presents index‐based features extraction technique (IBFET) to detect code clones at a very large‐scale level to billions of LOC at file level granularity. We performed preprocessing, indexing, and clone detection for more than 324 billion of LOC using a Hadoop distributed environment, which is quite faster and more efficient as compared to existing distributed indexing and clone detection techniques; meanwhile, it detects all three types of clones efficiently. The MapReduce rule of divide and conquer is used for a count and retrieve the similar features between different systems. We evaluated the execution time, scalability, precision, and recall of IBFET by using a well‐known clone detection data set IJaDataset and BigCloneBench; furthermore, we compared the results with other state‐of‐the‐art tools. Our approach is faster, flexible, scalable, and provides accurate results with high authenticity and can be implemented at a large‐scale level. Junaid Akram, Majid Mumtaz, Ping Luo 0004 |
Softw. Pract. Exp. | 3 |
| 2019 | VCIPR: Vulnerable Code is Identifiable When a Patch is Released (Hacker's Perspective)abstractVulnerable source code fragments remain unfixed for many years and they always propagate to other systems. Unfortunately, this happens often, when patch files are not propagated to all vulnerable code clones. An unpatched bug is a critical security problem, which should be detected and repaired as early as possible. In this paper, we present VCIPR, a scalable system for vulnerability detection in unpatched source code. We present a unique way, that uses a fast, token-based approach to detect vulnerabilities at function level granularity. This approach is language independent, which supports multiple programming languages including Java, C/C++, JavaScript. VCIPR detects most common repair patterns in patch files for the vulnerability code evaluation. We build fingerprint index of top critical CVE's source code, which were retrieved from a reliable source. Then we detect unpatched (vulnerable/non-vulnerable) code fragments in common open source software with high accuracy. A comparison with the state-of-the-art tools proves the effectiveness, efficiency and scalability of our approach. Furthermore, this paper shows that how the hackers can easily identify the vulnerable software whenever a patch file is released. Junaid Akram, Ping Luo 0004 |
ICST | 3 |
| 2019 | An Integrated Software Vulnerability Discovery Model based on Artificial Neural NetworkabstractQuantitative approaches for software security are needed for effective testing, maintenance and risk assessment of software systems.Vulnerabilities that are present in a software system after its release represent a great risk.Vulnerability discovery models (VDMs) have been proposed to model vulnerability discovery and have has been fined to vulnerability data against calendar time.Though, these models have various shortcomings include changes and development of VDMs for different dataset due to diverse approaches and assumptions in their analytical formulation.There is a clear need for an intensive investigation on these models to enhance predictive accuracy of existing VDMs and adopt the actual behavior of software vulnerabilities which were not modeled previously.This study proposed an integrated model to predict a number of software vulnerabilities by hybridizing the Multi-Layer Perceptron (MLP) artifical neural network and Vulnerability Discovery Models.The proposed model is also widely applicable across various vulnerability datasets and models due to its input diversity by providing improved fitting and predictive accuracy.Further, the experimental results show that this model not only retained the properties of traditional parametric VDM models as well as MLP's good nonlinear mapping ability and useful generalization. Gul Jabeen, Ping Luo 0004, Junaid Akram, Akber Aman Shah |
SEKE | 2 |
| 2019 | An improved software reliability prediction model by using high precision error iterative analysis methodabstractSummary Software reliability deals with the probability that software will not cause the failure of a system in a specified time interval. Software reliability growth models (SRGMs) are used to predict future behaviour from known characteristics of software, like historical failures. With the increasing demand to deliver quality software, more accurate SRGMs are required to estimate the software release time and cost of the testing effort. Software failure predictions at early phases also provide an opportunity for investing in proper quality assurance and upfront resource planning. Up till now, many parametric software reliability growth models (PSRGMs) have been proposed. However, several limitations of them mean that their predictive capacities differ from one dataset to others. In this paper, to enhance the prediction accuracy of existing PSRGMs, a high precision error iterative analysis method (HPEIAM) has been proposed based on the residual errors. In HPEIAM, residual errors from the estimated results of SRGMs are considered as another source of data that can combine the residual error modification with artificial neural network sign estimator. The repeated computation of residual errors by SRGMs improves and corrects the prediction accuracy up to the expected level. The performance of HPEIAM is tested with several PSRGMs using two sets of real software failure data based on three performance criteria. Moreover, we have compared the estimated failures predicted by HPEIAM with genetic algorithm (GA)‐based prediction improvement. The results demonstrate that HPEIAM gives an improvement in goodness‐of‐fit and predictive performance for every PSRGM in initial few iterations. Gul Jabeen, Ping Luo 0004, Wasif Afzal |
Softw. Test. Verification Reliab. | 2 |
| 2018 | DroidCC: A Scalable Clone Detection Approach for Android Applications to Detect Similarity at Source Code LevelabstractAndroid became more popular and widely used operating system. It has been noticed that the code clones in Android apps make it difficult to maintain the security flaws in source code. To avoid these problems, it is essential to find, identify, evaluate and recover those code clones as early as possible. In this paper, we propose and design DroidCC, a novel clone detection approach in Android applications, that helps to detect different types of clones from APK's source code. A prototype has been developed and implemented on the dataset of almost 30,000 top rated Android apps. DroidCC detects type-1, type-2 and type-3 clones in Android apps at the source code level. It also detects the similar code fragments, that were injected into many applications, which might be an indication of spreading malware. Meanwhile it can detect full and partial level similarity between applications. We evaluate DroidCC clone detection approach on real time data-set and count the Recall and Precision, which is quite significant. Furthermore, our results show that our approach is very efficient and effective in detecting different types of clones to check the similarity level in Android applications. Junaid Akram, Zhendong Shi, Majid Mumtaz, Ping Luo 0004 |
COMPSAC (1) | 4 |
| 2018 | DCCD: An Efficient and Scalable Distributed Code Clone Detection Technique for Big CodeabstractCode clone detection is a very hot topic in the field of software maintenance, reuseability and security.There is still a lack of techniques to detect near-miss clones at different level of granularities, especially in big code.This paper presents Distributed Code Clone Detection (DCCD) technique, which detects clones from big code bases based on feature extraction.We performed preprocessing, indexing and clone detection for almost 27 TB of source code (324 billion LOC), DCCD is quite faster and efficient as compared to existing distributed indexing and clone detection techniques, i.e. 36 times faster than Benjamin technique, which is 86 times faster than CCFinder.These two techniques are also distributed and just detect Type-1 and Type-2 clones, but our technique DCCD even detects Type-3 clones, efficiently.Our approach is faster, flexible, scalable and provides 87% accurate results with authenticity, ease of accessibility, upgradeability and maintainability. Junaid Akram, Zhendong Shi, Majid Mumtaz, Ping Luo 0004 |
SEKE | 4 |
| 2018 | A Unified Measurement Solution of Software Trustworthiness Based on Social-to-Software Framework
Gul Jabeen, Ping Luo 0004, Xiaoling Zhu, Mei-Hua Liu |
J. Comput. Sci. Technol. | 3 |