EDBT 2026 Demo / reviewers in the wild / expert
Yutong Zeng
dblp:328/5623
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring security posture of NAS third-party packages ecosystem: an empirical analysis
Jianbin Xu, Yutong Zeng, Tao Leng, Pin Yang |
Autom. Softw. Eng. | 3 |
| 2026 | Snake in the Grass: A Hybrid Detection Method Targeting Malicious PyPI PackagesabstractAs an increasing number of reusable packages are available in software development, package ecosystems are becoming more mature. Python is one of the most popular programming languages today. PyPI, as the primary repository of Python packages, serves a critical role in Python software development. While PyPI enables registered users to publish open-source packages, attackers can disguise themselves as de velopers to distribute malicious packages. We present a hybrid malicious python package detection framework named PCPD, which combines static anomaly detection and dynamic running analysis to detect lurking malicious packages. Using PCPD, we analyze a real-world dataset containing 414,880 PyPI packages and successfully narrow the inspection scope to 324 suspicious packages, reducing the manual review workload for repository managers by 99.48%. We have identified 121 packages with malicious behavior using PCPD and made all samples publicly available for security researchers to help mitigate these risks and protect the community. Additionally, we discover potential factors that enable malicious packages to evade scrutiny from repositories and associated threat signals. Cheng Huang 0003, Yutong Zeng, Genpei Liang, Yutong Du |
IEEE Trans. Reliab. | 2 |
| 2025 | From Soup to Nuts: A Hierarchical Relation-Based IP Attribution Approach for Cyber Threat TracebackabstractIP attribution plays a crucial role in the domain of network measurement and cyber threat intelligence, as it enables the traceback of malicious activities and the analysis of IP-Autonomous System Number (ASN) relation. This process provides critical evidence and intelligence for cyber law enforcement. However, current IP attribution method, Whois, face significant challenges because of IPv4 exhaustion and the privacy regulations. Consequently, an efficient and adaptable approach is needed to address these limitations. In this paper, we propose Scaling-RotatE, a novel IP attribution method based on knowledge graph link prediction, designed to infer the ASN ownership of IP addresses. Firstly, we extract multi-dimensional asset features and construct a IP-ontology knowledge graph. To enhance feature embedding, we introduce an innovative strategy with learnable modulus, embedding entities and relationships in the complex vector space to capture varying relationship strengths. A self-adversarial negative sampling technique is employed to enhance the model's generalization capability. Finally, leveraging knowledge graph link prediction techniques, we uncover latent IP-ASN relations, enabling more accurate IP attribution. Experimental results demonstrate that our model achieves a HITS@10 score of 0.831, outperforming numbers of previous methods and establishing a new state-of-the-art performance in IP attribution. Yutong Zeng, Tao Leng, Cheng Huang 0003 |
HPCC | 2 |
| 2025 | Wolf in Sheep's Clothing: Shearing the Camouflage of Malicious Java Components in MavenabstractIn recent years, software supply chain attacks have become increasingly prevalent, prompting considerable research into detecting malicious packages within relevant repositories. With the popularity bolstered by the widespread adoption of open-source practices, Java become one of the preferred languages among modern developers. However, the issue of malware detection in Java components remains unresolved. Most prior approaches suffer from insufficient code coverage and coarse-grained representation, making them unsuitable for Java components.In this paper, we propose an innovative solution calledSheartailored for detecting malicious Java components.Shearfirstly analyzes all methods in the component and locates potential malicious code snippets based on sensitive calls, as slice-level analysis provides a better understanding of the specific malicious activities. Secondly, statements depending on sensitive call sites are extracted and embedded into vectors for further detection instead of function-level representation which is coarse-grained facing the dynamic features in Java. The corresponding experimental results show thatSheareffectively identifies the malicious semantics hidden in the code slices by leveraging the neural network model, outperforming currently available tools to a great extent. Through real-world validation,Sheardetected 51 components with malicious characteristics out of 68,273, demonstrating its practical feasibility. This study introduces the first Java malicious component detection method suitable for real-world scenarios, carrying considerable practical significance in bolstering defenses within the software supply chain. Yutong Zeng, Cheng Huang 0003, Jiaxuan Han, Genpei Liang, Shuyi Jiang |
IEEE Trans. Software Eng. | 1 |
| 2025 | Towards Secure Code Generation With LLMs: A Study on Common Weakness EnumerationabstractAutomated code generation has revolutionized software development, enabling developers to accelerate project timelines and reduce manual coding errors significantly. As reliance on these technologies grows, the inherent weaknesses of generated code become increasingly apparent. Recent studies have shown that code produced by AI is not inherently safer or of higher quality than human-written code, often replicating existing vulnerabilities.To this end, we propose SECURECODER, which integrates Retrieval-Augmented Generation (RAG) with Common Weakness Enumeration (CWE). SECURECODER first utilizes the advanced reasoning capabilities of large language models (LLMs) to generate natural language descriptions of the code’s core business logic and functionality. Then, from a semantic perspective, it matches the requirements of the code generation task with the CWE descriptions through a multi-label classification process. Finally, based on the matched CWE, SECURECODER generates a list of security guidelines the code generation model must adhere to. Breaking down end-to-end code generation tasks into single-target tasks that LLMs excel at ensures that the generated code not only meets functional requirements but also adheres to best security practices, thereby enhancing the interpretability of the automated code generation process. After evaluating 2 programming languages and 7 LLMs on Coploit-generated code, SECURECODER has great generalization capability and could be applied to more programming languages and vulnerability types. SECURECODER could significantly decrease the security weakness in the AI-generated code and is able to mitigate more than 65% of vulnerabilities exposed to software developers. Compared to the baseline open-source LLMs, code vulnerabilities were reduced by at least 14% and the code business logic was not affected. Yuqiang Sun 0001, Cheng Huang 0003, YaoHui Guan, Yutong Zeng, Yang Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2024 | SGCML: Detecting Hacker Community Hidden in Chat GroupabstractHacker communities on various platforms have similar interests in sharing malware, vulnerability knowledge, and other illegal artifacts. Therefore, it is helpful to identify hacker communities to detect potential security issues. Many researchers have concentrated on detecting hacker communities on online social networks, such as Twitter and underground forums. However, the problem of detecting hacker communities in chat groups remains unresolved. To address two main challenges in detecting hacker communities in chat groups, this paper presented a method named Self-optimizing Graph Clustering based on Multitask Learning (SGCML). On the one hand, differing from traditional social networks, there are no direct edge relationships between users in chat groups, thereby this paper defines meta-paths as edges to reveal how hackers are connected. On the other hand, this paper presents a self-optimizing graph clustering method based on multi-task learning to address the issue of labeled data scarcity. According to the experiments, SGCML outperforms baselines to a great extent and can correctly identify significant nodes and structures in the hacker community, demonstrating its effectiveness in detecting hacker communities on chat groups. Tao Leng, Chang You, Yutong Zeng |
TrustCom | 5 |
| 2024 | GraySniffer: A Cliques Discovering Method for Illegal SIM Card Vendor Based on Multi-Source Data
Tao Leng, Chang You, Shuangchun Luo, Yutong Zeng |
TrustCom | 5 |
| 2022 | CSCD: A Cyber Security Community Detection Scheme on Online Social Networks
Yutong Zeng, Honghao Yu, Tiejun Wu, Xing Lan |
ICDF2C | 1 |
| 2022 | Research on Analog Integrated Circuit Test Parameter Set Reduction Based on XGBoost
Yindong Xiao, Yutong Zeng, Ke Liu 0005, Chong Hu |
J. Electron. Test. | 2 |