EDBT 2026 Demo / reviewers in the wild / expert
Liwei Zheng
dblp:04/4077
· DBLP profile ↗
22ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-7641-6369ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFHNN-DDI: Molecular fragment-based hypergraph neural networks for drug-drug interaction prediction
Liwei Zheng, Kaibiao Lin, Zhongqi Cai |
Neurocomputing | 1 |
| 2025 | Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsabstractIn recent years, large language models (LLMs) have made significant advancements in arithmetic reasoning.
However, the internal mechanism of how LLMs solve arithmetic problems remains unclear.
In this paper, we propose explaining arithmetic reasoning in LLMs using game-theoretic interactions.
Specifically, we disentangle the output score of the LLM into numerous interactions between the input words.
We quantify different types of interactions encoded by LLMs during forward propagation to explore the internal mechanism of LLMs for solving arithmetic problems.
We find that (1) the internal mechanism of LLMs for solving simple one-operator arithmetic problems is their capability to encode operand-operator interactions and high-order interactions from input samples.
Additionally, we find that LLMs with weak one-operator arithmetic capabilities focus more on background interactions.
(2) The internal mechanism of LLMs for solving relatively complex two-operator arithmetic problems is their capability to encode operator interactions and operand interactions from input samples.
(3) We explain the task-specific nature of the LoRA method from the perspective of interactions. Leilei Wen, Liwei Zheng, Zhihua Wei 0001, Wen Shen 0002 |
NeurIPS | 2 |
| 2025 | SCoVerLLM: Smart Contract Vulnerability Detection via LLM-Based In-Context and Chain-of-Thought PromptsabstractAs a key application of the blockchain technology, smart contracts have been adopted in various domains such as finance and the Internet of Things. However, their potential vulnerabilities can lead to significant economic losses, efficient and accurate vulnerability detection methods are essential to guarantee their security. Existing methods mostly rely on predefined rules or classification models, which suffer from high maintenance costs and limited semantic understanding about the smart contracts. To address this issue, this paper proposes SCoVerLLM (Smart Contract Vulnerability Detection via LLM-Based In-Context and Chain-of-Thought Prompts), which is designed to enhance the performance of smart contract vulnerability detection by using LLMs. SCoVerLLM combines prediction information generated by deep learning models with similar contract examples, and leverages In-Context Learning prompts and structured Chain-of-Thought templates to guide LLMs in step-by-step analyzing the logic of contracts for vulnerability detection. Experimental results show that SCoVerLLM outperforms existing four methods, including MANDO and Mythril, in terms of multiple metrics, with improvements of 10.72% to 19.20% in Accuracy, 8.70% to 18.51% in Precision, and 10.09% to 25.08% in F1. Xiguo Gu, Weili Xu, Zhanqi Cui, Liwei Zheng |
SMC | 5 |
| 2025 | SEOCD: Detecting obsolete code comments by fusing semantic features and expert features
Zhanqi Cui, Shifan Liu, Li Li 0114, Liwei Zheng |
Expert Syst. Appl. | 4 |
| 2024 | From Domain Models to Natural Language Requirement Documents: An Exploration of Requirement Missing Problems (P)abstractIn software engineering, precise requirement specifications are crucial, often derived from natural language documents prone to ambiguity and inconsistency.This can lead to critical omissions, necessitating costly maintenance.This study introduces RCADM, a method combining NLP and domain models to identify missing requirements in software documentation, enhancing the elicitation process. Tiankuo Wang, Liwei Zheng |
SEKE | 2 |
| 2024 | Combining Deep Learning and Expert Rules for Smart Contract Vulnerability DetectionabstractSmart contracts usually hold a large amount of digital assets, which can cause substantial losses if these contracts have vulnerabilities. Thus, it is essential to adequately detect possible vulnerabilities in smart contracts before deployment. There are many types of vulnerabilities in smart contracts, and different detection methods have their own unique advantages, some vulnerabilities may be more suitable for expert rule-based methods, while some vulnerabilities are more suitable for deep learning-based methods. A single detection method usually fails to fully use its ability to detect vulnerabilities. To address the above problems, we propose a composite approach named CDE-VD (Combining Deep Learning and Expert Rules for Smart Contract Vulnerability Detection) to improve the performance of vulnerability detection. The method divides smart contract samples into deep learning-prone sam-ples and expert rule-prone samples by classifying them before detection, and extracts expert rule features to train the smart contract detection method classifier to predict the category of the samples under analysis, then selects the suitable method for detection. The experimental results show that the vulnerability detection performance of CDE-VD outperforms that of single detection methods. Compared with the SOTA method MANDO, CDE-VD achieves average improvements of 3.22%, 2.32%, 9.25%, and 6.54% in terms of the Accuracy, Precision, Recall, and F1-score for five categories of vulnerabilities such as access control and time manipulation, respectively, which indicates that category prediction of the smart contract samples could improve vulnerability detection performance. Senlin Ren, Xiguo Gu, Liwei Zheng, Zhanqi Cui |
SMC | 4 |
| 2024 | TS-FL: Software Fault Localization Based on Teacher-Student NetworkabstractAutomated fault localization methods can expedite the process for developers to locate faulty code in complex software systems. Existing fault localization methods improve performance by combining the suspicious scores from different kinds of fault localization methods. Among these, the suspicious scores of mutation-based fault localization methods, commonly referred as mutation features, have been proven to effectively enhance fault localization performance. However, collecting mutation features requires generating a large number of mutants and executing test cases for each mutant, which demands sub-stantial computational resources and time. Additionally, certain code statements lack mutation features because no mutant can be generated for them, which affect the performance of fault localization. To address this, this paper proposes a Teacher and Student network-based Fault Lecalization (TS-FL) method. Firstly, a BiLSTM-based classifier is used to extract the deep semantic features of code statements, and the suspicious scores calculated by spectrum-based and mutation-based fault localization methods are used as the spectrum features and mutation features of the code statements, respectively. Then, a teacher-student network is constructed, and a mutual learning strategy is used to collaboratively train the teacher and student network, enabling the student network to learn the mutation feature information from the teacher network and thereby enhance its fault localization performance. The experimental results on Defects4J show that, without using mutation features, TS-FL can locate 36, 36, and 35 more faulty statements than spectrum-based fault localization methods Ochiai, Tarantula, and DStar, and can locate 8 more faulty statements than deep learning-based fault localization method TRANSFER-FL, in terms of Top-1. Jiale Zhang 0002, Liwei Zheng, Zhanqi Cui |
SMC | 3 |
| 2024 | A Fast Crash Reproduction Method for Android Applications Based on Widget Hierarchy GraphsabstractTo improve the efficiency of fixing bugs, mobile application developers must reproduce bugs reported by testers or users as quickly as possible. In some cases, automated testing tools can help developers reproduce crashes. However, these tools were not designed for reproducing bug reports. They are not efficient at reproducing crashes. To focus testing resources on suspicious widgets, we propose CrPDroid, a fast crash reproduction method for Android applications based on widget hierarchy graphs. First, it builds a widget hierarchy graph by using the project file of the application under test; then, it locates suspicious widgets by analyzing the bug report and the project file of the application under test and calculates the fitness of each widget according to the widget hierarchy graph; finally, it uses the fitness of widgets to guide automated testing to reproduce crashes quickly. To evaluate the effectiveness of CrPDroid, experiments are conducted on real Android application bug reports, and the crash reproduction tool ReCDroid, ReproBot and automated testing tools APE and PUMA are compared with CrPDroid. The experimental results show that CrPDroid successfully reproduces 15 bug reports that cause Android app crashes. In addition, compared with APE, PUMA, ReCDroid and ReproBot, the average time for CrPDroid to reproduce crashes decreased by 76.87%, 81.94%, 95.58% and 76.55%, and the total number of operations on suspicious widgets in the same period of testing time increased by 44.07%, 87.57%, 88.70% and 68.93% on average, respectively. Zhanqi Cui, Gaoyi Lin, Liwei Zheng |
IEEE Internet Things J. | 3 |
| 2024 | CrossFuzz: Cross-contract fuzzing for smart contract vulnerability detectionabstractSmart contracts are computer programs that run on a blockchain. As the functions implemented by smart contracts become increasingly complex, the number of cross-contract interactions within them also rises. Consequently, the combinatorial explosion of transaction sequences poses a significant challenge for smart contract security vulnerability detection. Existing static analysis-based methods for detecting cross-contract vulnerabilities suffer from high false-positive rates and cannot generate test cases, while fuzz testing-based methods exhibit low code coverage and may not accurately detect security vulnerabilities. The goal of this paper is to address the above limitations and efficiently detect cross-contract vulnerabilities. To achieve this goal, we present CrossFuzz, a fuzz testing-based method for detecting cross-contract vulnerabilities. First, CrossFuzz generates parameters of constructors by tracing data propagation paths. Then, it collects inter-contract data flow information. Finally, CrossFuzz optimizes mutation strategies for transaction sequences based on inter-contract data flow information to improve the performance of fuzz testing. We implemented CrossFuzz, which is an extension of ConFuzzius, and conducted experiments on a real-world dataset containing 396 smart contracts. The results show that CrossFuzz outperforms xFuzz, a fuzz testing-based tool optimized for cross-contract vulnerability detection, with a 10.58% increase in bytecode coverage. Furthermore, CrossFuzz detects 1.82 times more security vulnerabilities than ConFuzzius. Our method utilizes data flow information to optimize mutation strategies. It significantly improves the efficiency of fuzz testing for detecting cross-contract vulnerabilities. Huiwen Yang, Xiguo Gu, Xiang Chen 0005, Liwei Zheng, Zhanqi Cui |
Sci. Comput. Program. | 4 |
| 2023 | Software Fault Localization Based on Combining Information Retrieval and Mutation AnalysisabstractInformation Retrieval-based Bug Localization (IRBL) and Mutation-based Fault Localization (MBFL) are two widely used static and dynamic fault localization techniques, respectively. IRBL takes less time and utilizes more static information of software, while MBFL achieves high accuracy and the results are not easily affected by coincidental correctness test cases. However, the granularity of IRBL is coarse and MBFL consumes a lot of time to generate and execute mutants. In this paper, we propose IRMBFL (Information Retrieval and Mutation Analysis Based Software Fault Localization), a software fault localization technique that combines information retrieval and mutation analysis. First, the suspiciousness of source code files is measured by calculating the text similarity between the bug report and the source code to extract the files which may contain bugs. Then, the extracted files are mutated and tested. Finally, the bug statements are located by analyzing the changes in the execution results of the test cases. The experiments are conducted on the Defects4J dataset and$E_{inspect}{@} n$and EXAM are used as evaluation metrics to evaluate the performance of IRMBFL. The experimental results show that IRMBFL locates 14 and 3 more bug statements than BugLocator and Metallaxis for$E_{inspect}{@}n$when$n=1$. IRMBFL outperforms BugLocator on all projects and outperforms Met-allaxis on 2 out of 6 projects in terms of EXAM. In addition, the average bug localization time overhead of IRMBFL is reduced from 73.87% to 99.78% than Metallaxis. Liwei Zheng, Li Li 0114, Zhanqi Cui |
ATS | 3 |
| 2023 | TBCUP: A Transformer-based Code Comments Updating Approach
Shifan Liu, Zhanqi Cui, Xiang Chen 0005, Li Li 0114, Liwei Zheng |
COMPSAC | 6 |
| 2023 | AFL2oop: Loop Coverage Guided Greybox Fuzz TestingabstractFuzz testing automatically generates and executes test cases, to detect more defects by covering more logical and state spaces of the program under test (PUT).However, it becomes more difficult to adequately test the PUT with increasing size and code complexity.Studies have shown that complex code is more likely to contain defects, and the loop is one of the main reasons for increased code complexity.Therefore, it is necessary to thoroughly test the loops, but existing fuzzers cannot focus on the loops of the PUT.To address this issue, we design a loop interval coverage metric to measure the testing adequacy of the loop.Additionally, we propose a greybox fuzz testing approach named AFL 2 oop (AFL for Loop), which uses loop coverage as guidance.First, we analyze the loops of the PUT and expand the bitmap.Then, fuzz testing is guided by loop interval coverage and branch coverage.A prototype tool is implemented based on the proposed method, and experiments are carried out on four real-world software programs, such as LibXml2, LibMing, etc.The results show that AFL 2 oop achieves higher coverage, triggers more crashes, and reproduces defects faster than AFL and FairFuzz. Haochen Jin, Liwei Zheng, Zhanqi Cui |
SEKE | 2 |
| 2023 | CIDFuzz: Fuzz testing for continuous integrationabstractAbstract As agile software development and extreme programing have become increasingly popular, continuous integration (CI) has become a widely used collaborative work method. However, it is common to make changes frequently to a project during CI. If existing testing methods are applied to CI directly, it will be difficult to make testing resources focus on changes generated by CI, which results in insufficient testing for changes. To solve this problem, we propose a fuzz testing method for CI. First, differential analysis is performed to determine the change points generated during CI, change points are added to the taint source set, and static analysis is conducted to calculate the distances between each basic block and the taint sources. Then, the project under test is instrumented according to the distances. During fuzz testing, testing resources are allocated based on seed coverage to test the change points effectively. Using the proposed methods, we implement CIDFuzz as a prototype tool, and experiments are conducted on four open‐source projects that use CI. Experimental results show that, compared with AFL and AFLGo, CIDFuzz can reduce the time costs of covering change points up to 39.59% and 41.64%, respectively. Also, CIDFuzz can reduce the time costs of reproducing vulnerabilities up to 34.78% and 25.55%. Jiaming Zhang 0008, Zhanqi Cui, Xiang Chen 0005, Huiwen Yang, Liwei Zheng, Jianbin Liu |
IET Softw. | 5 |
| 2023 | OC-Detector: Detecting Smart Contract Vulnerabilities Based on Clustering Opcode InstructionsabstractSmart contracts are programs running on blockchain. In recent years, due to the persistent occurrence of security-related accidents in smart contracts, the effective detection of vulnerabilities in smart contracts has received extensive attention from researchers and engineers. Machine learning-based vulnerability detection techniques have the advantage that they do not need expert rules for determining vulnerabilities. However, existing approaches cannot identify vulnerabilities when the versions of smart contract compilers are updated. In this paper, we propose OC-Detector (Opcode Clustering Detector), a smart contract vulnerability detection approach based on clustering opcode instructions. OC-Detector learns the characteristics of opcode instructions to cluster them and replaces opcode instructions belonging to the same cluster with the ID of the cluster. After that, the similarity between the contract under analysis and contracts in the vulnerability database is calculated to identify vulnerabilities. The experimental results demonstrate that OC-Detector improves the F1 value of detecting vulnerabilities from 0.04 to 0.40 compared to DC-Hunter, Securify, SmartCheck and Osiris. Additionally, compared to DC-Hunter, the F1 value is improved by 0.27 when detecting vulnerabilities in smart contracts compiled by different versions of compilers. Xiguo Gu, Liwei Zheng, Huiwen Yang, Shifan Liu, Zhanqi Cui |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2022 | DeltaFuzz: Historical Version Information Guided Fuzz Testing
Jiaming Zhang 0008, Zhanqi Cui, Xiang Chen 0005, Huanhuan Wu, Liwei Zheng, Jian-Bin Liu |
J. Comput. Sci. Technol. | 5 |
| 2013 | A Consensus Model for Multiple Attribute Group Decision Making Problems Based on Interval Fuzzy NumberabstractSometimes, we find decision situations in which it is difficult to express preferences by means of concrete preference degrees. The consistency measure is a vital basis for consensus model of group decision making, and includes two sub problems: individual consistency measure and consensus measure. In this paper, a consensus model for multiple attribute group decision making problems in which the experts use interval fuzzy number to represent their attribute values is presented and a two-stage (a consensus process and a selection process) approach is proposed to solve it. Zhiying Lv, Tianmin Huang, Liwei Zheng |
SMC | 3 |
| 2012 | Modeling and Analyzing the Reliability and Cost of Service Composition in the IoT: A Probabilistic ApproachabstractRecently, many efforts have been devoted to explore the integration of Internet of Things (IoT) and Service-Oriented Computing (SOC). These works allow the real-world devices to provide their functionality as web services. However, two important issues, unreliable service providing and resource constraints, make the modeling and analysis of service composition in IoT a big challenge. In this paper, we propose a probabilistic approach to formally describe and analyze the reliability and cost-related properties of the service composition in IoT. First, a service composition in IoT is modeled as a finite state machine (FSM) which focuses on the functional part. Then, we extend this FSM model to a Markov Decision Process (MDP), which can specify the reliability of service operations. Furthermore, we extend MDP with cost structure, which can represent the different service quality attributes for each operation, such as energy consumption, communication cost, etc. The desirable quality properties of the service composition are specified by a probabilistic extension of temporal logic PCTL. We adopt a well-established probabilistic model checker PRISM to verify and analyze those properties of our service composition models. Lixing Li, Zhi Jin 0001, Ge Li 0001, Liwei Zheng |
ICWS | 4 |
| 2010 | Automatic selection of print-worthy content for enhanced web page printing experienceabstractThe user experience of printing web pages has not been very good. Web pages typically contain contents that are not print-worthy or informative such as side bars, footers, headers, advertisements, and auxiliary information for further browsing. Since the inclusion of such contents degrades the web printing experience, we have developed a tool that first selects the main part of the web page automatically and then allows users to make adjustments. In this paper, we describe the algorithm for selecting the main content automatically during the first pass. The web page is first segmented into several coherent areas or blocks using our web page segmentation method that clusters content based on the affinity values between basic elements. The relative importance values for the segmented blocks are computed using various features and the main content is extracted based on the constraint of one DOM (Document Object Model) sub-tree and high important scores. We evaluated our algorithm on 65 web pages and computed the accuracy based on area of overlap between the ground truth and the extracted result of the algorithm. Suk Hwan Lim, Liwei Zheng, Jianming Jin, Huiman Hou, Jian Fan, Jerry Liu |
ACM Symposium on Document Engineering | 2 |
| 2010 | Modeling in agent oriented internetware frameworkabstractResearches on Internetware have gained daily expanding attentions and interests. Internetware intends to be a framework of Web-based software development. This paper we model an Agent oriented Internetware Framework. Four principles are followed when the modeling approach is developed. They are the autonomy principle, the abstract principle, the explicitness principle and the competence principle. For conducting the Agent-oriented Internetware computing, three types of agents are needed based on these principles. They are the capability providing agents, the task planning agents and the task request agents. Different types of agents have different responsibilities in the computing framework. Base on the framework, we define the assignment problem in the collaboration of the autonomous Internetware entities, and show that its complexity is NP-complete. Then we model it as a Kripke structure with normative systems on it. This, which will help us to discuss some absorbing issues like the robustness or applying power indices on it in the future, builds the bridge between our work and researches on normative systems, games, mechanisms, etc. An approach based on negotiation is given to solve the problem. Liwei Zheng |
Internetware | 1 |
| 2009 | Acronym extraction and disambiguation in large-scale organizational web pagesabstractIn this paper, we focus on the automatic extraction and disambiguation of acronyms in large-scale organizational web pages, which is important but difficult due to the diversity of acronyms and the scale of organizational web pages. We propose two novel algorithms to address the key problems in acronym extraction and disambiguation: (1) An unsupervised ranking algorithm to automatically filter out the incorrect acronym-expansion pairs. Different from the existing approaches, our method does not require any hand-crafted rules; (2) A graph-based algorithm to disambiguate ambiguous acronyms, which leverages the hyperlinks of pages to facilitate the acronym disambiguation. We evaluate the proposed approaches using two large-scale, real-world datasets in two different domains. Our experimental results show that our approach is domain independent, and achieves higher precision and recall than the existing methods. Shicong Feng, Yuhong Xiong, Conglei Yao, Liwei Zheng |
CIKM | 4 |
| 2009 | OfCourse: web content discovery, classification and information extraction for online course materialsabstractIn this paper we present OfCourse, a vertical search engine for online course materials. These materials have the following characteristics: they are scattered very sparsely in the university Web sites; and are generated by the teachers with totally different HMTL templates and layouts. These characteristics impose some challenges for Web Classification (to identify the course materials) and Web Information Extraction (to extract course metadata, such as course title, time and ID) from the identified course homepages. Here, we describe our proposed method to tackle these challenges, and the features of this system. OfCourse, containing over 60,000 courses from the top 50 universities in the US, is currently available for public access at http://fusion.hpl.hp.com/OfCourse/. Yuhong Xiong, Ping Luo 0001, Shicong Feng, Baoyao Zhou, Liwei Zheng |
CIKM | 7 |
| 2009 | Aggregation of autonomous Internetware entitiesabstractResearches on Internetware have gained daily expanding attentions and interests. Internetware intends to be a framework of Web-based software development. A key issue in this framework is how these Internetware entities aggregate to form a coalition to fulfill the newly occurred requirements. This paper assumes that the Internetware entities distributed in Internet are autonomous and builds a mechanism for the aggregation of autonomous Internetware entities driven by requirements. A function ontology has been constructed for allowing the requirements to be understandable by these Interenetware entities so that these entities can recognize the requirements and realize that they are able to contribute for the realization of the requirements. After that, the requester and the contributors will negotiate to form an effective coalition. Liwei Zheng |
Internetware | 2 |