Yutong Cheng

dblp:324/7354 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-8641-8079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models
abstract
Textual descriptions in cyber threat intelligence (CTI) reports, such as security articles and news, are rich sources of knowledge about cyber threats, crucial for organizations to stay informed about the rapidly evolving threat landscape. However, current CTI knowledge extraction methods lack flexibility and generalizability, often resulting in inaccurate and incomplete knowledge extraction. Syntax parsing relies on fixed rules and dictionaries, while model fine-tuning requires large annotated datasets, making both paradigms challenging to adapt to new threats and ontologies. To bridge the gap, we propose CTINexus, a novel framework leveraging optimized in-context learning (ICL) of large language models (LLMs) for data-efficient CTI knowledge extraction and high-quality cybersecurity knowledge graph (CSKG) construction. Unlike existing methods, CTINexus requires neither extensive data nor parameter tuning and can adapt to various ontologies with minimal annotated examples. This is achieved through: (1) a carefully designed automatic prompt construction strategy with optimal demonstration retrieval for extracting a wide range of cybersecurity entities and relations; (2) a hierarchical entity alignment technique that canonicalizes the extracted knowledge and removes redundancy; (3) an long-distance relation prediction technique to further complete the CSKG with missing links. Our extensive evaluations using 150 real-world CTI reports collected from 10 platforms demonstrate that CTINexus significantly outperforms existing methods in constructing accurate and complete CSKG, highlighting its potential to transform CTI analysis with an efficient and adaptable solution for the dynamic threat landscape.
Yutong Cheng, Osama Bajaber, Saimon Amanuel Tsegai, Dawn Song, Peng Gao 0008
EuroS&P1
2024 Characterizing, Detecting, and Correcting Comment Errors in Smart Contract Functions
abstract
NatSpec comments play an essential role in smart contracts. Their clear and informative format helps users gain an accurate understanding of smart contract functions and diminish financial risk. However, widespread non-adherence to NatSpec standards currently causes confusion for both end-users and developers. Current research often neglects the importance of NatSpec formats or solely emphasizes user-centric comments in smart contract generation. This oversight can hinder contract trustworthiness, code reusability, maintenance efficiency, and ultimately, the development of the community ecosystem. To bridge this gap, this paper presents the first empirical study on 253 verified contracts encompassing 16,620 functions from Etherscan, uncovering that 87 % of the smart contract functions have Comment Errors (CE) and pinpointing prevalent deviation patterns. Based on our findings, we propose CETerminator, an automated approach for detecting and rectifying CE in smart contract functions. Due to the scarcity of NatSpec-compliant comments for collected smart contract functions, CETerminator employs in-context learning on a large language model to generate NatSpec comments. The approach then compares the original and the generated comments, utilizing corpus-driven heuristic rules to identify and correct diverse error categories in the original comments. In our evaluation, CETerminator demonstrates a high token overlap rate for addressing missing comments. In addition, the average precision, recall, and F1-scores for handling inconsistency comments are 85.28 %, 86.48 %, and 85.85%, respectively, outperforming the baseline by 39.79%, 39.53%, and 39.84%.
Yutong Cheng, Zhengda Li
SSE1
2023 Hue: A User-Adaptive Parser for Hybrid Logs
abstract
Log parsing, which extracts log templates from semi-structured logs and produces structured logs, is the first and the most critical step in automated log analysis. While existing log parsers have achieved decent results, they suffer from two major limitations by design. First, they do not natively support hybrid logs that consist of both single-line logs and multi-line logs (Java Exception and Hadoop Counters). Second, they fall short in integrating domain knowledge in parsing, making it hard to identify ambiguous tokens in logs. This paper defines a new research problem, hybrid log parsing, as a superset of traditional log parsing tasks, and proposes Hue, the first attempt for hybrid log parsing via a user-adaptive manner. Specifically, Hue converts each log message to a sequence of special wildcards using a key casting table and determines the log types via line aggregating and pattern extracting. In addition, Hue can effectively utilize user feedback via a novel merge-reject strategy, making it possible to quickly adapt to complex and changing log templates. We evaluated Hue on three hybrid log datasets and sixteen widely-used single-line log datasets (Loghub). The results show that Hue achieves an average grouping accuracy of 0.845 on hybrid logs, which largely outperforms the best results (0.563 on average) obtained by existing parsers. Hue also exhibits SOTA performance on single-line log datasets.
Junjielong Xu, Qiuai Fu, Zhouruixing Zhu, Yutong Cheng, Zhijing Li 0007, Yuchi Ma, Pinjia He
ESEC/SIGSOFT FSE4