EDBT 2026 Demo / reviewers in the wild / expert
Wei Li 0254
dblp:64/6025-254
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-9374-4409ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-authorSecurity and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CodeMark: Contextual and Natural Watermarking for Tracing Code Snippet ProvenanceabstractDetermining the origins of code snippets has gained increasing attention due to the popularity of large language models and the concern about their misuse in generating unlicensed or malicious code. Watermarking is considered a working solution for tracing code snippet provenance. However, source code watermarking requires more stringent and intricate rules than natural language or software watermarking, since one needs to ensure both readability and functionality of the watermarked code snippets. To this end, we propose a novel watermarking systemCodeMark, featured by variable renaming as the key. Surrounding variable renaming, several challenges emerge such as determining renaming candidates, defining the variable context, and providing diverse variable substitutes, etc. The design ofCodeMarkconquers these challenges by an end-to-end learning system, which interprets the code context through Graph Neural Networks (GNNs) and generates natural substitutes fitting the context by distilling from CodeBERT. Experiments illustrate thatCodeMarksurpasses the state-of-the-art watermarking systems in terms of watermarking requirements. Wei Li 0254, Borui Yang, Yujie Sun 0001, Suyu Chen, Yuting Chen 0001, Liyao Xiang |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsabstractThe expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or misuse of the code for malicious purposes. Hence it is important to track the ownership of source code, in which watermarking is a major technique. Yet, drastically different from natural languages, source code watermarking requires far stricter and more complicated rules to ensure the readability as well as the functionality of the source code. Hence we introduce SrcMarker, a watermarking system to unobtrusively encode ID bitstrings into source code, without affecting the usage and semantics of the code. To this end, SrcMarker performs transformations on an AST-based intermediate representation that enables unified transformations across different programming languages. The core of the system utilizes learning-based embedding and extraction modules to select rule-based transformations for watermarking. In addition, a novel feature-approximation technique is designed to tackle the inherent non-differentiability of rule selection, thus seamlessly integrating the rule-based transformations and learning-based networks into an interconnected system to enable end-to-end training. Extensive experiments demonstrate the superiority of SrcMarker over existing methods in various watermarking requirements. Borui Yang, Wei Li 0254, Liyao Xiang, Bo Li 0001 |
SP | 2 |
| 2024 | MiniTracker: Large-Scale Sensitive Information Tracking in Mini AppsabstractRunning on host mobile applications, mini apps have gained increasing popularity these days for its convenience in installation and usage. However, being easy to use allows mini apps to freely access a large amount of user information, mostly without close inspection of privacy violations. Hence it becomes a crucial issue to automatically track sensitive flows in mini apps. Although flow analysis has been widely studied, unique challenges emerge: the analysis tool should not only handle mini app-specific features such as flows that interweave between rendering and logic, and asynchronous executions, but also deal with problems raised by Javascript development: the performance tradeoff between precision and efficiency, and function aliases. To this end, we proposeMiniTracker, an automatic sensitive flow tracking tool which well handles mini app features, constructs assignment flow graphs as common representation across different host apps, searches function aliases, and analyzes the graph by property chains. We show our design choices achieve a sweet spot in the tradeoff between precision and efficiency, with superior performance compared to the state-of-the-art. We also perform a large-scale study on 150 k mini apps, which reveals the common leakage patterns and offers insights into the privacy threats of mini apps. Wei Li 0254, Borui Yang, Hangyu Ye, Liyao Xiang, Qingxiao Tao, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Learning Code-Query Interaction for Enhancing Code SearchesabstractCode search plays an important role in software development and maintenance. In recent years, deep learning (DL) has achieved a great success in this domain-several DL-based code search methods, such as DeepCS and UNIF, have been proposed for exploring deep, semantic correlations between code and queries; each method usually embeds source code and natural language queries into real vectors followed by computing their vector distances representing their semantic correlations. Meanwhile, deep learning-based code search still suffers from three main problems, i.e., the OOV (Out of Vocabulary) problem, the independent similarity matching problem, and the small training dataset problem. To tackle the above problems, we propose CQIL, a novel, deep learning-based code search method. CQIL learns code-query interactions and uses a CNN (Convolutional Neural Network) to compute semantic correlations between queries and code snippets. In particular, CQIL employs a hybrid representation to model code-query correlations, which solves the OOV problem. CQIL also deeply learns the code-query interaction for enhancing code searches, which solves the independent similarity matching and the small training dataset problems. We evaluate CQIL on two datasets (CODEnn and CosBench). The evaluation results show the strengths of CQIL-it achieves the MAP@1 values, 0.694 and 0.574, on CODEnn and CosBench, respectively. In particular, it outperforms DeepCS and UNIF, two state-of-the-art code search methods, by 13.6% and 18.1% in MRR, respectively, when the training dataset is insufficient. Wei Li 0254, Haozhe Qin, Shuhan Yan, Beijun Shen, Yuting Chen 0001 |
ICSME | 1 |
| 2019 | Reinforcement Learning of Code Search SessionsabstractSearching and reusing online code is a common activity in software development. Meanwhile, like many general-purposed searches, code search also faces the session search problem: in a code search session, the user needs to iteratively search for code snippets, exploring new code snippets that meet his/her needs and/or making some results highly ranked. This paper presents Cosoch, a reinforcement learning approach to session search of code documents (code snippets with textual explanations). Cosoch is aimed at generating a session that reveals user intentions, and correspondingly searching and reranking the resulting documents. More specifically, Cosoch casts a code search session into a Markov decision process, in which rewards measuring the relevances between the queries and the resulting code documents guide the whole session search. We have built a dataset, say CosoBe, from StackOverflow, containing 103 code search sessions with 378 pieces of user feedback. We have also evaluated Cosoch on CosoBe. The evaluation results show that Cosoch achieves an average NDCG@3 score of 0.7379, outperforming StackOverflow by 21.3%. Wei Li 0254, Shuhan Yan, Beijun Shen, Yuting Chen 0001 |
APSEC | 1 |
| 2019 | CocoQa: Question Answering for Coding Conventions Over Knowledge GraphsabstractCoding convention plays an important role in guaranteeing software quality. However, coding conventions are usually informally presented and inconvenient for programmers to use. In this paper, we present CocoQa, a system that answers programmer's questions about coding conventions. CocoQa answers questions by querying a knowledge graph for coding conventions. It employs 1) a subgraph matching algorithm that parses the question into a SPARQL query, and 2) a machine comprehension algorithm that uses an end-to-end neural network to detect answers from searched paragraphs. We have implemented CocoQa, and evaluated it on a coding convention QA dataset. The results show that CocoQa can answer questions about coding conventions precisely. In particular, CocoQa can achieve a precision of 82.92% and a recall of 91.10%. Repository: https://github.com/14dtj/CocoQa/ Video: https://youtu.be/VQaXi1WydAU. Tianjiao Du, Junming Cao, Qinyue Wu, Wei Li 0254, Beijun Shen, Yuting Chen 0001 |
ASE | 4 |
| 2019 | Constructing a Knowledge Base of Coding Conventions from Online ResourcesabstractCoding conventions are a set of coding guidelines used by software developers to improve the readability of source code, increase software maintainability, and promote the reuse of coding patterns.In this paper, we introduce CCBase, a knowledge base of coding conventions, that was constructed from online resources.Specifically, CCBase was constructed as follows.We designed the ontology of the coding convention domain, crawled data related to coding conventions from a variety of online resources, and then extracted entities and relations using an NLP-enabled rule matching method.To uncover the latent relations, we further proposed a similarity metric to reveal the similar-to and relate-to relations, and developed a RCE algorithm to establish a unified type hierarchy of coding conventions.The resulting knowledge base contains 3139 coding conventions for Java and C++, with 3761 entities and 767 relations.Furthermore, we have extended the usability of CCBase by developing a question answering system on the base.We have conducted experiments to evaluate CCBase.The experimental results show that CCBase has a wide coverage on entities and relations in coding conventions domain, and the QA system achieves an F1 score of 84.5% on 214 questions raised in StackOverflow. Junming Cao, Tianjiao Du, Beijun Shen, Wei Li 0254, Qinyue Wu, Yuting Chen 0001 |
SEKE | 4 |